One prompt set, many models, two ways to score.
- throughput
- Time to first token, tokens/sec, inter-chunk and total latency, token usage.
- image_gen
- Generation latency plus the saved PNGs, for side-by-side review.
One file, ~/.llmbench/config.yaml, holds every API
key and every model. Twenty providers are known by name.
- openai
- anthropic
- gemini
- moonshot
- deepseek
- xai
- groq
- mistral
- together
- fireworks
- openrouter
- perplexity
- cerebras
- qwen
- nvidia
- nebius
- deepinfra
- sambanova
- flux
Local, no key required
- ollama
- vllm
- lmstudio
- llamacpp
Published scores from four sources, refreshed daily. No API key needed
to read them, in the terminal or in the table below.
- huggingface
- Open LLM Leaderboard v2: IFEval, BBH, MATH, GPQA, MUSR, MMLU-PRO.
- lmarena
- LMArena ELO, from head-to-head human preference voting.
- aider
- Aider Polyglot: multi-language code-editing pass rate.
- bundled
- Snapshot shipped inside the package, so it works offline.