# Deploy with self-serve deployments

Self-serve deployments give you reserved inference capacity on Crusoe's optimized inference engine and managed infrastructure. You choose a base model (or a fine-tuned adapter) and a deployment configuration, and Crusoe handles engine selection, tuning, autoscaling, and rate limiting. You get predictable performance, dedicated throughput, no shared rate limits, and per-GPU-hour billing you control.

## Prerequisites

- Install the OpenAI Python client (`pip install openai httpx`), if you use the Python examples.
- For Low-Rank Adaptation (LoRA) adapters, bring a checkpoint from a successful training job completed with [Serverless Fine-Tuning](/serverless-fine-tuning/overview).

## 1. Log in or create an account

Log in to the [Crusoe Cloud Console](https://console.crusoecloud.com) or [Create an account](/create-an-account).

After you log in, switch to the **Intelligence Foundry** app in the bottom-left of the [console](https://console.crusoecloud.com/).

## 2. Generate an API key and authenticate

To create an API key through the [console](https://console.crusoecloud.com):

1. From the [console](https://console.crusoecloud.com), click **Admin** in the bottom-left corner.
2. Select **Security** > **[Intelligence API keys](https://console.crusoecloud.com/security/inference-api-keys)** from the left navigation.
3. Click **Create**.
4. (Optional) Enter an alias for your key.
5. (Optional) Enter an expiration date for your key.
6. Copy the **API key**. Make sure that you save the key in a secure location before leaving the page.

### Authenticate against the API

Use your Intelligence API key to authenticate across your intelligence API calls and self-serve deployment management API calls.

1. Export the token and base URL in your shell:

   ```shell
   export API_TOKEN='<your token>'
   export INFERENCE_URL='https://api.inference.crusoecloud.com/v1/chat/completions'
   export DEPLOYMENT_URL='https://api.crusoecloud.com/v1/projects/{project_id}/foundry/selfserve/'
   ```

2. For inference requests, construct an OpenAI client pointed at the Crusoe gateway. Every Python example on this page assumes you have this `openai_client` in scope:

   ```python
   from openai import OpenAI
   import httpx, os

   openai_client = OpenAI(
       api_key=os.environ["API_TOKEN"],
       url=f"os.environ['INFERENCE_URL']",
       http_client=httpx.Client(proxy=None, trust_env=False),
   )
   ```

The full OpenAPI specification is published at [api.intelligence.crusoecloud.com/docs](https://api.intelligence.crusoecloud.com/docs).

## 3. Choose a base model and deployment configuration

Define your deployment by selecting the base model you want to serve and the deployment configuration you want to optimize for.

### Deployment configurations

Each deployment offers one or more optimization profiles. Pick the configuration that matches your workload requirements and Crusoe will apply the corresponding engine, hardware, and optimizations for you. There's no hand-tuning required.

| Configuration      | Optimization                                                                                          | Best for                                                                   |
| ------------------ | ----------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------- |
| **Responsiveness** | Low latency, optimized for time-to-first-token                                                        | Interactive applications, real-time inference, latency-sensitive workloads |
| **Throughput**     | Cost efficiency at scale, optimized for token volume                                                  | Batch processing, high-volume workflows, cost-per-token minimization       |
| **Balanced**       | Hybrid blend of throughput and responsiveness, optimized to support moderate token volume and latency | General purpose production traffic                                         |

### Supported models

Refer to [Available models](/self-serve-deployments/available-models) for the full list of supported base models, which you can also deploy with LoRA adapters trained through [Serverless Fine-Tuning](/serverless-fine-tuning/overview).

## 4. Create a deployment

Create a deployment from the console or API to get started.

**UI:**

1. Sign in to the [console](https://console.crusoecloud.com/) and switch to the **Intelligence Foundry** app in the bottom-left corner.

2. Select **[Self-Serve Deployments](https://console.crusoecloud.com/foundry/deployments)** from the **Inference** section of the left navigation. The page lists every deployment in your project with its status, model, hardware, replicas, and metadata.

3. Click **Create deployment** and complete the form:

   - Select a base model and, optionally, a fine-tuned checkpoint
   - Select a deployment configuration (Responsiveness, Throughput, or Balanced)
   - Select the number of replicas you want your deployment to support

   The hourly cost associated with your chosen deployment configuration is displayed before you confirm.
**cURL:**

```bash
curl "$DEPLOYMENT_URL/deployments" \
  -X POST \
  -H "Authorization: Bearer $API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "flavor_id": "<flavor_id>",
    "deployment_name": "<deployment_name>",
    "replicas": <replica_count>
  }'
```

After the deployment creation is initiated, your new deployment will appear in the self-serve deployment table. Click any row in the self-serve deployment table to view activity logs, endpoint metadata, and its current replica count.

## 5. Check deployment status

A new deployment provisions reserved capacity, which can take up to 40 minutes to complete. A deployment moves through the following states during its lifecycle:

| State          | Description                                                  |
| -------------- | ------------------------------------------------------------ |
| `Creating`     | Capacity is being provisioned and the engine is starting up. |
| `Ready`        | The deployment is ready to serve traffic.                    |
| `Scaling up`   | The deployment is scaling up its active replicas.            |
| `Scaling down` | The deployment is scaling down its active replicas.          |
| `Syncing`      | The deployment alias is being updated.                       |
| `Failed`       | The deployment couldn't be created or updated.               |
| `Deleting`     | The deployment is being torn down.                           |
| `Deleted`      | The deployment has been removed.                             |

Check the status of your deployment from the deployment table on the **Self-Serve Deployments** page or directly through the API. The status updates to `Ready` when the deployment is available to serve traffic.

```bash
curl "$DEPLOYMENT_URL/deployments" \
  -X GET \
  -H "Authorization: Bearer $API_TOKEN" \
```

## 6. Run inference

When the deployment is in `Ready` state, send requests to it through the OpenAI-compatible Chat Completions API using your API key. Pass the deployment alias as the `model`.

**python:**

```python
import os
from openai import OpenAI

client = OpenAI(
    base_url=os.environ['INFERENCE_URL'],
    api_key=os.environ['API_TOKEN'],
)

response = client.chat.completions.create(
    model='<deployment alias>',
    messages=[
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Summarize the theory of relativity in one sentence."}
    ],
)

print(response.to_json())
```
**typescript:**

```typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: process.env.INFERENCE_URL,
  apiKey: process.env.API_TOKEN,
});

client.chat.completions
  .create({
    model: "<deployment alias>",
    messages: [
      { role: "system", content: "You are a helpful assistant." },
      {
        role: "user",
        content: "Summarize the theory of relativity in one sentence.",
      },
    ],
  })
  .then((response) => console.log(response));
```
**cURL:**

```shell
curl $INFERENCE_URL \
  --request 'POST' \
  --header 'Content-Type: application/json' \
  --header 'Accept: text/event-stream' \
  --header "Authorization: Bearer $API_TOKEN" \
  --data '{
    "model": "<deployment alias>",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Summarize the theory of relativity in one sentence."}
    ]
  }'
```

Because your deployment runs on reserved capacity, its throughput is bounded by replica count rather than a shared rate limit. To add headroom for traffic spikes, increase the replica count on the deployment in the next step.

## 7. Manage deployments

You can update a running deployment (for example, to change replica count or the deployment alias) or delete one you no longer need. When you delete a deployment, billing stops.

To view all deployment management options, select the three-dot icon on any deployment row on the **Self-Serve Deployments** page. The menu exposes options to edit the deployment alias, update the replica count, or delete the deployment.

### Configure notifications

By default, [notifications](/managed-ai/notifications) are sent when a self-serve deployment is created, deleted, or scaled. You can manage your preferences from the console's [Notifications settings](https://console.crusoecloud.com/notifications) page.

### Edit a deployment alias

To update the alias for a deployment:

**UI:**

1. Define a unique name for your deployment.
2. Confirm your new deployment name.
**cURL:**

```bash
curl "$DEPLOYMENT_URL/deployments/{id}" \
  -X PATCH \
  -H "Authorization: Bearer $API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "deployment_name": "<deployment_name>"
  }'
```

The deployment status updates to `Syncing` while the alias is updated, and returns to `Ready` when the update is complete.

### Update replica counts

To adjust the number of replicas for a deployment:

**UI:**

1. Select a value that reflects your expected traffic load and fits within your allotted quota.
2. Confirm your new replica count.
**cURL:**

```bash
curl "$DEPLOYMENT_URL/deployments/{id}" \
  -X PATCH \
  -H "Authorization: Bearer $API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "replicas": "<replica_count>"
  }'
```

The deployment status updates to `Scaling up` or `Scaling down` based on the direction of the change, and returns to `Ready` when the update is complete. The deployment can still serve traffic while the replica count is being adjusted.

### Delete a deployment

After you confirm deletion, the deployment status updates to `Deleting`, and the deployment disappears from the list once the deletion is complete. Billing stops when the deployment is deleted.

```bash
curl "$DEPLOYMENT_URL/deployments/{id}" \
  -X DELETE \
  -H "Authorization: Bearer $API_TOKEN" \
  -H "Content-Type: application/json" \
```

### View deployment details

To view the deployment overview, activity log, metadata, and sample code for inferencing, select the deployment alias from the **Self-Serve Deployments** page or use the API.

```bash
curl "$DEPLOYMENT_URL/deployments/{id}" \
  -X GET \
  -H "Authorization: Bearer $API_TOKEN"
```

## 8. (Optional) Deploy a fine-tuned model

Self-serve works with Crusoe serverless fine-tuning to help you get your tuned models to production with one click. A LoRA adapter you train there is registered in the same model registry as the base models, so it appears in the models list as soon as training completes.

You have three ways to deploy a fine-tuned checkpoint:

- **From the Self-Serve Deployments page:** Follow the [Create a deployment](#4-create-a-deployment) steps, then select your fine-tuned checkpoint from the list after you select the corresponding base model architecture.
- **From a Fine-tuned model's [Jobs](https://console.crusoecloud.com/foundry/fine-tuning/jobs) page:** Click the three-dot menu next to the checkpoint you want to deploy and select **Deploy**.
- **From the self-serve deployments API:** First, retrieve your `fine_tuned_model` identifier for your desired fine-tuned model checkpoint and include that in your self-serve deployment creation request.

  ```bash
  curl "$DEPLOYMENT_URL/deployments" \
    -X POST \
    -H "Authorization: Bearer $API_TOKEN" \
    -H "Content-Type: application/json" \
    -d '{
      "fine_tuned_model_id": "<fine_tuned_model_identifier>",
      "flavor_id": "<flavor_id>",
      "deployment_name": "<deployment_name>",
      "replicas": <replica_count>
    }'
  ```

Because fine-tuning and deployment share the same registry and API conventions, you can iterate quickly: train a new adapter, point a new deployment at it, and shift traffic—without rebuilding your serving stack.

## Next steps

- Need optimization beyond the standard configurations? [Contact us](https://www.crusoe.ai/contact-sales) about Tailored Deployments
- For an overview of Crusoe's Managed AI options, see [Managed AI](/managed-ai/overview)