/ Live
Full live dashboard
Per-endpoint detail, reliability heatmaps, and per-sample variance plots will live here. Below: the live tiles and trend chart from the home page.
Endpoint
Status
tok/s
FORGE · gemma4-26b
RTX 5090 · vLLM · 47m ago · vllm_metrics
OK
348.1
HYDRA-R · GPU 0
AMD Radeon Pro R9700 · llama.cpp · 47m ago · llamacpp_timings
OK
75.6
HYDRA-R · GPU 1
AMD Radeon Pro R9700 · llama.cpp · 47m ago · llamacpp_timings
OK
75.2
NEST · GPU 0
RTX 3060 12 GB · vLLM · 47m ago · vllm_metrics
OK
38.2
NEST · GPU 1
RTX 3060 12 GB · vLLM · 46m ago · vllm_metrics
OK
37.3
Ollama Cloud · gemma4-cloud
cloud passthrough via Jester · 46m ago · server_wall
OK
85.9
Ollama Cloud · glm-5.2
cloud passthrough via Jester · 46m ago · server_wall
OK
104.8
SCOUT · gemma4-26b
RTX 3090 Ti + 3090 · vLLM TP=2 · 46m ago · vllm_metrics
OK
250.5
TITAN · Engine A
2× RTX 3090 · vLLM TP=2 · 46m ago · vllm_metrics
OK
235.9
TITAN · Engine B
2× RTX 3090 · vLLM TP=2 · 46m ago · vllm_metrics
OK
242.7
Decode tok/s · 24H trend
Each point is one sample, taken at the top of the hour: one warmup run discarded, one timed run recorded. Same prompt every time. When an hour has no successful run, the line dives to the floor and a red dot marks the incident — timeout, rate-limit, or other non-OK status. We don't smooth incidents into the curve. Full methodology.
More charts coming as we add features. Reliability heatmap and per-sample variance scatter are next.