OZZZER · AI NEWS2 of 3 free stories opened
← Back to AI News

Coding & Development · 1 Oct 2026 · 20:31 CEST

Benchmarking LLM Inference at Scale with AIPerf

NVIDIA · 1 Oct 2026 · 20:31 CESTRead original at NVIDIA ↗
Share
LinkedInXFacebookWhatsApp
Benchmarking LLM Inference at Scale with AIPerf

Publisher preview · OZZZER analysis pending editorial review.

PUBLISHER ARTICLE PREVIEW

From the original article

You’re deploying a model on a system. It starts up, prompts are getting responses. Now the hard question: Is this fast?

Your instincts might lead you to send curl commands, hand-roll an asyncio script, or vibe code yet another one-off load generator. All of these paths have the same problem: single-process performance limits, Python’s GIL capping concurrency, or numbers measured against a reference you built yourself. Either way, you end up with results you can’t fully trust, attached to tooling you’ll have to rewrite the moment requirements change.

What you need is a load client that can saturate a real server without becoming the bottleneck, produce output you can act on, and take five minutes to configure, not five hours. That’s NVIDIA AIPerf.

AIPerf is the designated successor to GenAI-Perf and is a ground-up rewrite. The design choices reflect hard lessons from running LLM benchmarks at scale:

For this walkthrough we’ll use Qwen3-0.6B served through vLLM. The model choice is deliberate; it’s small enough to run on a single GPU and fast enough to iterate on without waiting. The point isn’t to benchmark Qwen3-0.6B specifically; it’s to establish the measurement loop. Once you have that, swapping in a different model or endpoint is a one-flag change.

One platform note: on aarch64, the crick dependency ships as source-only and requires a C toolchain (build-essential on Debian/Ubuntu, Development Tools on RHEL). If the install stalls on that package, that’s why.

With the server up and AIPerf installed, we can now run our first profile:

A few flags here are doing more work than they look like:

--synthetic-input-tokens-stddev 0 and --output-tokens-stddev 0 pin the workload to exactly 128 input and 128 output tokens per request. This reproduces a commonly used static benchmark that holds request and output lengths constant.

--extra-inputs min_tokens:128 and --extra-inputs ignore_eos:true tell the model to actually emit 128 tokens rather than stopping early. Without these, the output token count is a suggestion. The model stops whenever it naturally finishes, which can be well short of your target OSL. Throughput numbers end up lower than they should be, and they’re not reproducible across runs.

--streaming is not optional if you want to

Source

NVIDIA · 1 Oct 2026 · 20:31 CEST

Open the original at NVIDIA ↗