Skip to main content

7 docs tagged with "inference"

View all tags

Available models

Base models supported by Serverless Inference on Crusoe's Intelligence Foundry

Managed AI

Choose a Crusoe Managed AI product to run, adapt, or deploy AI models without managing GPU infrastructure

Metrics

Learn which metrics Serverless Inference records for your models and how to query them with PromQL

Self-Serve Deployments

Reserved-capacity inference deployments on Crusoe's managed infrastructure with predictable performance and no shared rate limits

Serverless Inference

Choose how to run inference on hosted models based on your traffic patterns and latency requirements

Serverless rate limits

How Serverless Inference enforces tokens-per-minute and requests-per-minute limits, and how to interpret 429 and 503 responses

Usage and billing

Usage and billing information for Serverless Inference, Serverless Fine-Tuning, and Self-Serve Deployments