mirror of
https://github.com/magnus919/agent-skills.git
synced 2026-09-11 19:47:12 +03:00
docs(litellm): harden lifecycle and rate-limit guidance
Authorized durable follow-up squash merge. Required validate passed at exact head 42e62196bbdca341848e584985e687b0ec56ebf4; advisory droid-review failed due BYOK ApplyPatch tooling error with no review findings.
This commit is contained in:
@@ -97,10 +97,29 @@ Semantics that matter:
|
||||
`POST /team/new` (with `members_with_roles`, limits), `/team/info`,
|
||||
`/team/member_add`, `/team/update`; `POST /user/new`, `GET /user/info`. Roles:
|
||||
PROXY_ADMIN, PROXY_ADMIN_VIEW_ONLY, ORG_ADMIN (EE), INTERNAL_USER,
|
||||
INTERNAL_USER_VIEW_ONLY, TEAM, CUSTOMER. Model access groups
|
||||
INTERNAL_USER_VIEW_ONLY, TEAM, CUSTOMER. API bodies use the lowercase enum
|
||||
literals, for example `user_role: "internal_user"`. An uppercase role name in
|
||||
`/user/new` returned a Pydantic `literal_error` 422 in a v1.98.0 observation on
|
||||
2026-08-25. Treat that as a version-and-date-scoped observation, not a promise
|
||||
about every release.
|
||||
|
||||
In that same v1.98.0 observation, `POST /user/new` returned a newly minted API
|
||||
key. Decide whether the key is needed before creating the user. If it is not
|
||||
needed, revoke or delete it through the documented API after confirming the
|
||||
exact key, user, and intended scope. A key that remains in the database is a
|
||||
persisted credential record, not an "untracked credential"; record its owner and
|
||||
lifecycle state so it can be audited. Model access groups
|
||||
(`model_info.access_groups`) let keys/teams be granted a group name instead of
|
||||
enumerated models.
|
||||
|
||||
Before an authorized update, deletion, or cleanup, confirm the target identifier,
|
||||
the intended scope (key, user, team, or deployment), and a rollback path. Read
|
||||
back the current values first, record them, and preserve them for restoration.
|
||||
Prefer block/revoke when temporary containment is sufficient; use deletion only
|
||||
when retention and recovery requirements permit it. This reference has no live
|
||||
v1.98.0 service: the enum and auto-mint statements above are source-scoped
|
||||
observations, not a reproduction performed here.
|
||||
|
||||
## Spend tracking
|
||||
|
||||
- Every request writes a spend log row (tokens, cost, model, key hash, end user);
|
||||
|
||||
@@ -85,13 +85,60 @@ secret and restarting; the regenerate flow re-encrypts stored credentials under
|
||||
unusable key and bricks the deployment. Symptom of salt-key trouble at startup:
|
||||
`Error decrypting value`.
|
||||
|
||||
### D) 429 rate limits — three different sources
|
||||
### D) 429 rate limits — identify the limiter and boundary
|
||||
|
||||
Read the wording: key/team rpm/tpm rejections come before any provider call and
|
||||
name the limit; budget exhaustion raises Budget/TokenBudget errors naming spend vs
|
||||
max (check `GET /key/info`); provider 429 carries `<Provider>Exception`, gets retried
|
||||
with backoff, then cools the deployment down (mode A). Remember budgets fail open
|
||||
without a DB and rate limits don't apply to admins.
|
||||
A 429 is not necessarily upstream. Classify the response and logs before changing
|
||||
configuration: a provider 429 includes `<Provider>Exception`; a proxy-side limit
|
||||
may name a key or team limiter; a router cooldown can report that no deployment is
|
||||
available. Budget exhaustion is a separate Budget/TokenBudget error. Budgets can
|
||||
fail open without a DB, and rate-limit checks may not apply to proxy admins, so use
|
||||
an internal-user test key when verifying enforcement.
|
||||
|
||||
LiteLLM has multiple limiter scopes. Key-level `tpm_limit`, `rpm_limit`, and
|
||||
`max_parallel_requests` apply to that key; per-model key limits use
|
||||
`model_tpm_limit` and `model_rpm_limit`. Team limits apply to team membership,
|
||||
while `general_settings.global_max_parallel_requests` is proxy-wide. A model or
|
||||
deployment `rpm`/`tpm` value in `litellm_params` may guide weighted routing rather
|
||||
than enforce a hard ceiling unless `enforce_model_rate_limits` is enabled. Do not
|
||||
infer a global, model, or deployment limit from a key-level message, and treat any
|
||||
future limiter or version-specific implementation as unverified until checked in
|
||||
the pinned release.
|
||||
|
||||
A bounded v1.98.0 observation from 2026-08-25 included `Limit type: tokens`, a
|
||||
`Current limit: 400000`, remaining tokens, and a reset timestamp. In that sample,
|
||||
LiteLLM's `ProxyRateLimitError` was raised by
|
||||
`proxy/hooks/parallel_request_limiter_v3.py` before the provider call, with logging
|
||||
through `common_request_processing.py _handle_llm_api_exception`. These are
|
||||
implementation details of that observation, not universal behavior, and no live
|
||||
v1.98.0 reproduction was performed for this reference.
|
||||
|
||||
To investigate a suspected key limiter, first capture the complete response,
|
||||
request ID, key identity without logging the secret, selected model/deployment,
|
||||
and the configured key/team/model values. Verify with the same key through the
|
||||
proxy: readiness, `/v1/models`, `/health?model=<name>`, response headers, and one
|
||||
small representative request. Prefer these bounded API and gateway checks before
|
||||
any authorized database inspection. If a change is approved, confirm the exact
|
||||
key/team/deployment target, intended scope, and rollback owner; read and record
|
||||
prior `tpm_limit`, `rpm_limit`, `max_parallel_requests`, and related values before
|
||||
changing them. Preserve those values and restore them after the test. Use an
|
||||
explicit, complete payload, for example:
|
||||
|
||||
```json
|
||||
{
|
||||
"key": "<token>",
|
||||
"tpm_limit": null,
|
||||
"rpm_limit": null,
|
||||
"max_parallel_requests": null
|
||||
}
|
||||
```
|
||||
|
||||
Send that payload to `POST /key/update` only after the confirmation gate. Explicit
|
||||
`null` requests clearing these key fields; omitting a field leaves its prior value
|
||||
in place. Clearing key limits cannot prove that a team, global, model, deployment,
|
||||
or provider limiter is absent. Re-read the effective configuration and restore the
|
||||
recorded values through the same authorized API. Inspect Postgres directly only
|
||||
when the API/gateway evidence is insufficient and the operator has approved the
|
||||
specific read scope.
|
||||
|
||||
### E) ContextWindowExceededError
|
||||
|
||||
|
||||
Reference in New Issue
Block a user