Know exactly how your AI performs under load.

InferGauge benchmarks LLM endpoints with live dashboards, reproducible benchmarks, and actionable performance insights. Measure latency, TTFT, throughput, goodput, cost, and capacity from your own machine.

Getting started

One command to install.

curl -fsSL https://raw.githubusercontent.com/Nexus-InferGauge/infergauge-releases/main/install.sh | sh

Or download the tarball manually

Not on Linux? Windows · macOS

Three commands to your first result.

Capabilities

Everything you need to benchmark AI inference.

Concurrency

Find the point where your application starts to slow down as user count increases.

Latency

Measure average, p95, and p99 latency under realistic workloads.

Time to First Token

Understand queueing delays before responses begin streaming.

Goodput

Track the percentage of requests that actually meet your SLOs.

Cost

Estimate token usage, cost per request, and projected spend.

Performance Score

Summarize every run with transparent metrics and reproducible results.

Built for developers.

Runs locally

Prompts, API keys, and benchmark results stay on your machine.

Reproducible

Every benchmark is a small YAML file that anyone on your team can rerun.

CI ready

Detect regressions, compare runs, and automate performance testing.

Simple workflow.

01

Configure your endpoint

Define your benchmark with a small YAML configuration.

02

Run a benchmark

Launch tests from the CLI against local or hosted models.

03

Understand the results

Explore live dashboards, compare runs, and export reports.

Start measuring without guesswork.

Benchmark your AI applications, understand their limits, and ship with confidence.