AI Inference Platform & GPU Operations Control Plane
GPU fleet telemetry · model routing · request queue · failure/failover tracking · cost analysis
1 LAB GPU
31 SIMULATED GPUs
Model Performance MEASURED — from stored request records
| Model | Backend | Requests | Error Rate |
p50 (ms) | p95 (ms) | p99 (ms) |
TTFT (ms) | Tokens/sec | Throughput |
Latency percentiles computed from individual request records in SQLite — not hardcoded.
Load Test Results MEASURED
Reproducible load test: scripts/run_load_test.py · raw JSON in results/
Failure Scenarios MEASURED — before/after metrics
Reproducible: scripts/run_failure_scenario.py · raw JSON in results/failure_scenarios.json
Estimated Inference Cost
Per-Model Cost
Computed from actual request tokens + representative GPU pricing (see README).
Data sources: LAB 1 real GPU (RTX 4090 config) ·
SIMULATED 31 synthetic fleet GPUs ·
MEASURED load tests & failure scenarios