AI Infra Bench
How much real AI infrastructure engineering work can frontier models solve?
17 tasks4 runs per task
Results
Each configuration is evaluated four times per task.
| Task details | Model | Effort | Agent | ||||
|---|---|---|---|---|---|---|---|
gpt-6-astra | xhigh | codex 0.153.4 | 64.7% ± 4.8% | 31.3 | 30.3 | $325.01 | |
gpt-6-astra | medium | codex 0.153.4 | 55.9% ± 5.9% | 22.3 | 21.3 | $129.31 | |
gpt-6-astra | high | codex 0.153.4 | 54.4% ± 5.6% | 24.7 | 23.7 | $191.36 | |
gpt-5.6-sol | xhigh | codex 0.153.4 | 53.7% ± 12.8% | 92.9 | 91.7 | $432.48 | |
gpt-6-astra | low | codex 0.153.4 | 50.0% ± 3.4% | 18.8 | 17.8 | $92.37 | |
gpt-5.6-sol | high | codex 0.153.4 | 48.5% ± 5.6% | 61.2 | 60.1 | $260.95 | |
gpt-5.6-sol | medium | codex 0.153.4 | 41.2% ± 8.3% | 35.8 | 34.8 | $105.78 | |
gpt-5.6-sol | low | codex 0.153.4 | 36.8% ± 11.1% | 23.5 | 22.5 | $57.35 |
Tasks
17 offline tasks with execution-based behavioral and e2e tests.
Bug fix
Anthropic Inline System Template
Adapt Anthropic inline system messages to chat-template ordering constraints.
Serving API
Bug fix
ASR Chunk Spacing
Preserve word boundaries when merging multi-chunk ASR output.
Serving API, Model implementation
Bug fix
Async KV Token Accounting
Preserve exact computed-token accounting across asynchronous external KV loads.
Scheduler, KV cache data movement
Bug fix
Async Spec Placeholder Discard
Prevent stale speculative frames from underflowing async scheduler placeholder accounting.
Scheduler
Bug fix
Concurrent Config Refresh
Tolerate transient configuration-file read failures during concurrent cache refreshes.
Model, Frontend API
Feature
CPU Offload Reset Inflight
Reset CPU-offload cache state without releasing blocks still used by asynchronous transfers.
KV cache data movement
Page 1 of 3. Showing tasks 1 to 6 of 17.
