As of: 2026-09-13
Repo: apache/tooling-llmao
Product design (concepts/policy): apache/rai-private → services/llmao/README.md
How to run/use this software: repo README.md
Ops: Infra p6/modules/llmao/README.md
Doc split
| Kind | Where |
|---|---|
| Product concepts (project-centered, budgets, credentials) | rai-private design |
| What the code does today (short) | Done below — detail is the code |
| Planned UX (layouts, workflows, phases) | UX backlog below |
| How to run locally / PATs against proxy | README (how-to) |
Implemented enough for local production-shaped use: asfquart OAuth; LiteLLMBackend + fail-fast team cache warm; models.yaml definitions; PAT UX (My Keys / Other Keys); Models page (supply-path redaction for non–site-admins); secrets as dual YAML / eyaml intent; system Postgres + prisma setup; offline tests/mock_backend.py.
Models page: definitions-first table (name, id, Available Yes/No, Request a key → /keys/new?model=). Context, license, and hosting live in Details; supply-path fields stay site-admin in the modal. Available is policy (always true until P5) and in service and in LiteLLM. Self-hosted with nothing Healthy or no deployment is No. Sort puts No last. Per-process state is in the /models admin drill-down.
models.yaml vs LiteLLM: admin definitions; the proxy does not include this file.
litellm.yaml general_settings.store_model_in_db: true (Puppet may also set the env). self_hosted is a required boolean. Commercial rows need a static api_base (fail-fast). Self-host litellm_params.model is hosted_vllm/<name> (not openai/). Standalone watches config.yaml; Puppet restarts on change. YAML edit ≠ live deployment until llmao POSTs /model/new again.
GPU fleet: fleet.hosts (IP → listen port); box JSON listen only; VllmServer + FleetDeployment. /models is signed-in (host:port admin). The green badge is Healthy and LiteLLM has that api_base. Self-host: /model/new when Healthy, /model/delete on Unhealthy or Stalled when a route exists (models.yaml template). Commercial: /model/new at llmao startup. LiteLLM /health cached on the deployment by the skew runner.
Public port resolution differs by provider. Vast exposes its container→public mapping through an API and resolves automatically. RunPod does not, so a host row takes an optional fourth element pinning the public port. RunPod also reassigns the port on every pod recreate, even when the pod lands on the same machine — so a pinned value is expected to change, not to be stable.
The proxy serves end to end. A PAT reaches three self-hosted models through llm.apache.org with TLS on 443, the LiteLLM port bound to loopback, tool calling and vision on the models that support them, and per-key attribution in the spend log.
| model | ctx | modality | tools | reasoning default | sustained |
|---|---|---|---|---|---|
gemma4-26b |
131,072 | text+vision | yes | off | ~128 tok/s |
qwen3.8-27b |
131,072 | text+vision | yes | on | ~46 tok/s |
qwen3-8b |
40,960 | text | yes | on | ~54 tok/s |
Sustained figures are long generations. Short prompts measure round-trip overhead and run roughly half these. Single-stream decode is memory-bandwidth bound, so the models sit closer together than their sizes suggest — the difference between them is concurrency, not per-request speed.
Fit validation: models.yaml declares vram_gb and disk_gb, and a box checks them against nvidia-smi and statvfs before pulling weights. Requirements sum across co-resident servers. Silent when nvidia-smi is absent — an unknown is not a failure.
Observed state: the /models drill-down shows measured KV cache against the served context window, scraped from vLLM's /metrics on the transition into Healthy. A --max-model-len above what the cache holds makes vLLM hang rather than error, which was previously only visible by reading a startup log.
Remaining: long vLLM boot, config revision on the box, pending-assignment table, box smoke, and nothing restarting a GPU box automatically.
Open policy still: who creates automation PATs (A RAI / B Chair-VP / C any PMC — code provisional C). See design §5.1.1.
Edge cases that bite operators (LiteLLM down, pagination, master key drift, prisma path, etc.) are tracked as work items, not product design.
Three bottlenecks stacked in front of the GPUs, each hiding the next. Recorded here because the order matters more than any individual value.
Apache ProxyTimeout killed any generation over 60 seconds. It looked
exactly like a model truncating. Now 1800s.
vLLM's max_num_batched_tokens defaults to 2048, which held every box
to 2–3 concurrent requests while 90% of each card's KV cache sat idle —
requests spent 4.3× longer queued than being processed. Raised, the 27B
went from 2 concurrent to 38.
The value is not portable between cards: 16384 suits the 96GB and 80GB boxes, and on a 24GB card it OOM'd during an FP8 matmul and killed the engine outright. That box runs 4096. It scales with headroom, not with the model.
KV cache was never the constraint. It is consumed by context, not by request count — small prompts cannot fill it however many you send, which is why an early load test proved nothing.
The order that matters: transport, then correctness, then scheduler, then memory, then hardware. We measured hardware first, concluded a 96GB card was over-provisioned, and were wrong.
Tool-call parsers are per model and the names mislead. qwen3-8b needs
hermes, not qwen3_xml, despite being a Qwen model — it emits
<tool_call> blocks wrapping JSON, the Hermes convention. With the wrong
parser the call arrives as text in content, tool_calls is absent, and
nothing reports an error. Verify by running each candidate against the
model's real output offline; vllm:tool_call_parser_invocations
distinguishes "the parser declined" from "the model never called".
Two scripts in bin/, with no endpoints committed — they discover from
LiteLLM's routing table at runtime. llmao-saturate is stdlib-only.
llmao-smoke reads litellm.master_key from config.yaml via PyYAML.
LLMAO_KEY overrides it; load that from a file
(read -r LLMAO_KEY < ~/.llmao-key && export LLMAO_KEY). A key typed on
the command line is stored in shell history and is visible to other users
via ps.
| what it answers | |
|---|---|
llmao-smoke |
does every model still work: completion, reasoning control, tool calling, vision, sustained output, large prompt. Exits non-zero, so it works in CI |
llmao-saturate |
where does a box start queueing, and is the limit the scheduler or memory — which decides whether a bigger card would help |
llmao-smoke --direct bypasses LiteLLM; the difference between the two runs
is the gateway's own overhead. Do not run the two scripts together — smoke
queueing behind a saturation ramp produces numbers that look like a
regression and are not.
Claude Code works against the gateway with environment variables alone;
LiteLLM exposes /v1/messages, so no shim is needed.
ANTHROPIC_BASE_URL=https://llm.apache.org
ANTHROPIC_AUTH_TOKEN=<PAT>
ANTHROPIC_MODEL=gemma4-26b
CLAUDE_CODE_MAX_CONTEXT_TOKENS=120000
Three things that are not obvious:
- Set the context limit below the real window. vLLM rejects a request
where prompt +
max_tokensexceeds it, and a client that assumes 200k will auto-compact too late. - Pick a model whose reasoning is off by default. A reasoning model emits nothing for a minute while it thinks; the client abandons the stream and retries, and the retries stack.
reasoning_effortvocabularies differ. Qwen3.8-27B accepts onlyxhigh,medium,lowand 400s onhigh, which several clients send by default. Gemma accepts all five.
Tool use through the Anthropic path is currently broken — see open issues.
Product intent for budgets, roles, and reports: design §6. Below is what to build in the UI, in order.
My Keys | Other Keys (PMC+) | Models | Projects | Reports?
Home = role-aware launchpad (not keys-only)
Projects is the second pillar (envelopes + people). Budgets are not a floating top-level “Budgets” app without project context.
- Projects list: projects I’m in / I administer; mini envelope % for steward projects.
- Project overview (
/projects/<name>or equivalent):- Money meters: People vs Automation spend split (display split OK if one team budget under the hood)
- Period label (e.g. monthly · reset date)
- Grantor of the dollar ceiling (v1: Free Tier on first cfg default; later RAI / Security / …)
- By-person spend this period (transparent to project members)
- Export CSV (steward+)
- Automation summary + link toward Other Keys for that project
- Empty: “Envelope appears when first key is minted” / trial copy when RAI defines defaults
- Flashes on any POST
- Home: purpose line; primary CTA keys
- ≥ ~90% near-limit callouts (keys or projects)
- Top ~3 personal key (or project) usages this period
- Steward strip: 2–3 administered projects with % used
- My Keys: project column → project overview; optional near-limit badge
- Collapsed “For scripts” base URL hint (not hero)
- Members table: cap / used / Edit cap dialog (empty = no cap; cannot exceed envelope)
- Optional: lower people/automation sub-caps if policy allows
- Cannot raise outer ceiling (no fake Increase button — Request/raise is RAI)
- Quart flashes on save/error; later PMC email (design audit)
- Superuser-only: set/raise project envelope(s), period, trial/free/allocated type
- Dual hard people vs automation limits when RAI/LiteLLM support exists
- Flashes; PMC notification when email exists
- Foundation roll-up (RAI)
- Commons fairness / high-usage distributions (careful labels)
- Self-hosted utilization (capacity; Infra signals as needed)
- My usage (committer)
- Not in open product: donated credit burn-down (RAI-private vendor deals)
- Capacity meters (TPM/RPM/parallel) separate from $
- Models: wire Available (and key
modelsallow-list) when team / envelope policy exists; column already stubbed
- ASF vocabulary only (project, person, purpose, envelope)
- Script-first; secret once
- Quart flashes for mutations
- Proxy-down = loud banner, not silent zeros
- Supply-path model fields: site admin only (already in Models v1)
- Trial/free default amounts and duration
- Hard dual budgets vs display-only people/automation split
- Narrow stewards later (Chair vs any PMC)?
- Capacity fair-share defaults from Infra
hosted_vllm/— done inmodels.yaml. New/model/newbodies use that prefix. LiteLLM DB rows created asopenai/stay until deleted and re-pushed.- Deployment registration is not idempotent. A host that changes
address leaves the old deployment behind, and nothing removes it — one
model had accumulated twelve identical rows, invisible in both UIs. The
bases - fleet_basesbranch of the skew check is exactly this case and should delete rather than only report. Pair with reconciliation at startup. A Models sync action when skew is found is a later UX. models.yamldoes not recordreasoning_effortlevels per model (Qwen 3.8 lists them; not a uniform field).- Harden PAT against LiteLLM pagination / delete ids
- Automation creator policy after RAI decides §5.1.1
- Site admin via
raiPMC (optional keep cfg list) - PMC notification email on key/budget lifecycle
- Advisor / richer routing
- Nothing restarts a GPU box. All three were launched by hand; a
host restart leaves the box up and the model down. RunPod pods carry
their launch command in
--docker-argsand recover; the hand-launched box does not.
Nothing is stored on the GPU boxes. They run vLLM and nothing else — no database, no application state. Prompt content exists in GPU memory for the duration of one request and is overwritten. Deleting a pod leaves nothing behind.
What is retained is on the gateway host: token counts, model, latency, and
a hashed key for attribution. LiteLLM's messages, response and
proxy_server_request columns can hold full request and completion text
and are switched off; every row contains {}.
That is a current state rather than a principle. Prompt history is useful — for debugging, evaluation, possibly retrieval — and we will likely want some form of it. When we do it should be a deliberate decision with retention limits and an access model, decided openly rather than arrived at by default.
Both providers are on secure tiers. A host operator with hardware access could in principle observe a request while it is processed; that is true of any rented hardware, is contractually prohibited on both platforms, and is a live window rather than an archive.
For the pilot the guidance is that anything sent through should be treated as visible to llmao admins: public and project-internal material, yes; credentials or embargoed work, no. That belongs on the Models page.
| Tree | Role |
|---|---|
rai-private services/llmao/ |
Product design |
| tooling-llmao | Software + STATUS (this file) |
p6 modules/llmao |
Production deploy (in production) |