﻿---
title: stack es ml start-trained-model-deployment cli command
description: Start a trained model deployment. Behaviour flags: --dry-run — validate all inputs and exit without performing any action 
url: https://www.elastic.co/elastic/docs-builder/docs/4075/reference/elastic-cli/cli/stack/es/ml/start-trained-model-deployment
applies_to:
  - Elastic Cloud Serverless: Preview
  - Elastic Stack: Preview
---

# stack es ml start-trained-model-deployment cli command
<cli-modifiers>
</cli-modifiers>

```bash
elastic stack es ml start-trained-model-deployment \
  --model-id <model-id> \
  [options]
```

Start a trained model deployment.
**Behaviour flags:**
`--dry-run` — validate all inputs and exit without performing any action

## Options

<definitions>
  <definition term="--model-id string required">
    The unique identifier of the trained model. Currently, only PyTorch models are supported.
  </definition>
  <definition term="--cache-size string">
    The inference cache size (in memory outside the JVM heap) per node for the model.
    The default value is the same size as the `model_size_bytes`. To disable the cache,
    `0b` can be provided.
  </definition>
  <definition term="--deployment-id string">
    A unique identifier for the deployment of the model.
  </definition>
  <definition term="--number-of-allocations number">
    The number of model allocations on each node where the model is deployed.
    All allocations on a node share the same copy of the model in memory but use
    a separate set of threads to evaluate the model.
    Increasing this value generally increases the throughput.
    If this setting is greater than the number of hardware threads
    it will automatically be changed to a value less than the number of hardware threads.
    If adaptive_allocations is enabled, do not set this value, because it’s automatically set.
  </definition>
  <definition term="--priority enum">
    The deployment priority
    **Values:** normal, low
  </definition>
  <definition term="--queue-capacity number">
    Specifies the number of inference requests that are allowed in the queue. After the number of requests exceeds
    this value, new requests are rejected with a 429 error.
  </definition>
  <definition term="--threads-per-allocation number">
    Sets the number of threads used by each model allocation during inference. This generally increases
    the inference speed. The inference process is a compute-bound process; any number
    greater than the number of available hardware threads on the machine does not increase the
    inference speed. If this setting is greater than the number of hardware threads
    it will automatically be changed to a value less than the number of hardware threads.
  </definition>
  <definition term="--timeout string">
    Specifies the amount of time to wait for the model to deploy.
  </definition>
  <definition term="--wait-for enum">
    Specifies the allocation status to wait for before returning.
    **Values:** started, starting, fully_allocated
  </definition>
  <definition term="--adaptive-allocations string">
    Adaptive allocations configuration. When enabled, the number of allocations
    is set based on the current load.
    If adaptive_allocations is enabled, do not set the number of allocations manually.
  </definition>
  <definition term="--error-trace">
    When set to `true` Elasticsearch will include the full stack trace of errors
    when they occur.
  </definition>
  <definition term="--filter-path string">
    Comma-separated list of filters in dot notation which reduce the response
    returned by Elasticsearch.
    **Repeatable:** pass `--filter-path` multiple times to supply more than one value
  </definition>
  <definition term="--human">
    When set to `true` will return statistics in a format suitable for humans.
    For example `"exists_time": "1h"` for humans and
    `"exists_time_in_millis": 3600000` for computers. When disabled the human
    readable values will be omitted. This makes sense for responses being consumed
    only by machines.
  </definition>
  <definition term="--pretty">
    If set to `true` the returned JSON will be "pretty-formatted". Only use
    this option for debugging only.
  </definition>
  <definition term="--input-file string">
    path to a JSON file to use as command input
  </definition>
  <definition term="--dry-run">
    validate all inputs and exit without performing any action (preview changes without applying them)
  </definition>
</definitions>


## Global Options

<definitions>
  <definition term="--json">
    output as JSON
  </definition>
</definitions>