> ## Documentation Index
> Fetch the complete documentation index at: https://vastai-80aa3a82-auto-openapi-update-8483c1e9.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Update Endpoint

> Updates autoscaling configuration on an existing endpoint; endpoint must not be managed by a deployment.



## OpenAPI

````yaml /api-reference/openapi.yaml put /api/v0/endptjobs/{id}/
openapi: 3.1.0
info:
  title: Vast.ai API
  description: >-
    Vast.ai REST API for managing GPU cloud instances, machine operations, and
    AI/ML workflows.


    ## AI Agent Quick-Start


    Install the CLI skill for your agent (Claude Code, Cursor, Windsurf, etc.):
      npx skills add vast-ai/vast-cli

    CLI reference:
    https://raw.githubusercontent.com/vast-ai/vast-cli/master/vastai/SKILL.md

    SDK reference:
    https://raw.githubusercontent.com/vast-ai/vast-cli/master/vastai_sdk/SKILL.md


    ## Auth

    All endpoints require `Authorization: Bearer $VAST_API_KEY`.

    Get your key at: https://cloud.vast.ai/manage-keys/


    ## Key Quirks

    - `gpu_ram` in CLI = GB; in REST API = MB (CLI auto-converts)

    - SSH keys must be registered BEFORE creating an instance (VM: no recovery;
    Docker: can add post-create)

    - `onstart` field is limited to 4048 characters -- gzip+base64 for longer
    scripts

    - `POST /api/v0/asks/{id}/` (create instance) returns `new_contract` as the
    instance ID, not `id`

    - Poll trap: if `actual_status` becomes `exited`, `unknown`, or `offline` it
    will never reach `running` -- destroy and retry
  version: 1.0.0
  contact:
    name: Vast.ai Support
    url: https://discord.gg/vast
servers:
  - url: https://console.vast.ai
    description: Production server
security:
  - BearerAuth: []
paths:
  /api/v0/endptjobs/{id}/:
    put:
      tags:
        - Serverless
      summary: Update Endpoint
      description: >-
        Updates autoscaling configuration on an existing endpoint; endpoint must
        not be managed by a deployment.
      parameters:
        - name: id
          in: path
          required: true
          schema:
            type: integer
          description: ID of the endpoint group to update
      requestBody:
        content:
          application/json:
            schema:
              type: object
              properties:
                min_load:
                  type: number
                  description: Minimum floor load in perf units/s (token/s for LLMs)
                  example: 0
                min_cold_load:
                  type: number
                  description: Updated minimum cold load threshold.
                  example: 0.05
                target_util:
                  type: number
                  description: Target capacity utilization (fraction, max 1.0)
                  example: 0.9
                cold_mult:
                  type: number
                  description: >-
                    Cold/stopped instance capacity target as multiple of hot
                    capacity target
                  example: 2.5
                cold_workers:
                  type: integer
                  description: Min number of workers to keep 'cold' when you have no load
                  example: 5
                max_workers:
                  type: integer
                  description: Max number of workers your endpoint group can have
                  example: 20
                max_queue_time:
                  type: number
                  description: Updated maximum acceptable queue time in seconds.
                  example: 90
                target_queue_time:
                  type: number
                  description: >-
                    Updated desired queue time in seconds; must be <=
                    max_queue_time.
                  example: 15
                endpoint_name:
                  type: string
                  description: Deployment endpoint name
                  example: my_endpoint
                autoscaler_instance:
                  type: string
                  description: Autoscaler deployment environment.
                  example: prod
                endpoint_state:
                  type: string
                  description: New activation state for the endpoint.
                  example: active
                  enum:
                    - active
                    - suspended
                    - stopped
                inactivity_timeout:
                  type: number
                  description: Updated seconds of inactivity before suspension.
                  example: 600
                overrecruit_ratio:
                  type: number
                  description: Updated over-recruit ratio for traffic spike absorption.
                  example: 1.3
      responses:
        '200':
          description: 'Returns {success: true}'
          content:
            application/json:
              schema:
                type: object
                properties:
                  success:
                    type: boolean
                    description: Always true on success.
              example:
                success: true
        '400':
          description: >-
            Endpoint is deployment-managed, or "Endpoint {endptjob_id} for user
            {user_id} not found."
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Error'
        '403':
          description: Host-only accounts cannot update endpoints or use serverless
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Error'
      security:
        - BearerAuth: []
components:
  schemas:
    Error:
      type: object
      properties:
        error:
          type: string
        msg:
          type: string
  securitySchemes:
    BearerAuth:
      type: http
      scheme: bearer
      description: API key must be provided in the Authorization header

````