Custom Models
Deploy and query your own models — finetunes and base-model deployments — behind your API key. One deployment per account; spend caps and scale-to-zero included.
Get Modal GPU/scaling options and live rates
Get Modal GPU/scaling options and live rates
Get all serving endpoints
Get all serving endpoints
Create a serving endpoint from a succeeded fine-tune job
Create a serving endpoint from a succeeded fine-tune job
Get an endpoint's status, invoke URL, and running cost
Get an endpoint's status, invoke URL, and running cost
Terminate the endpoint permanently
Terminate the endpoint permanently
Deploy the endpoint with the chosen compute options
Deploy the endpoint with the chosen compute options
Stop serving on this endpoint
Stop serving on this endpoint
Deploy a succeeded fine-tune as this endpoint's next version
Deploy a succeeded fine-tune as this endpoint's next version
Roll the deployment back to its previous version
Roll the deployment back to its previous version
Remove an old version from the rollback history
Remove an old version from the rollback history
Get the endpoint's usage log and latency/cost rollup
Get the endpoint's usage log and latency/cost rollup
Preview held-out validation samples from the training data
Preview held-out validation samples from the training data
Create an API key scoped to this endpoint only
Create an API key scoped to this endpoint only
Run forecast/generation inference on a tuned endpoint
Run forecast/generation inference on a tuned endpoint
Create embeddings with a tuned endpoint
Create embeddings with a tuned endpoint