A
AIPerf // CONFIG LAB
Workload specification studio
RESET
Module 01 / 04
Endpoint-Protokoll
Definiere Modell und API-Ziel für die spätere Messung.
Latency
Saubere Einzelanfragen
Throughput
Maximale parallele Last
Request Rate
Kontrollierte Poisson-QPS
Goodput
SLO-konforme Leistung
ShareGPT
Reale Multi-Turn-Chats
Trace Replay
Zeitstempelgetreue Wiedergabe
Server URL(s)
Erforderlich
http://localhost:8000
Eine URL pro Zeile; mehrere Ziele werden Round-Robin angesprochen.
Modellkatalog
Erforderlich
Nemotron 3.5 Lightning 30B A3B — nvidia/nemotron-3.5-lightning-30b-a3b
Nemotron 3 Ultra 550B A55B — nvidia/nemotron-3-ultra-550b-a55b
Nemotron 3 Super 120B A12B — nvidia/nemotron-3-super-120b-a12b
Nemotron 3 Nano 30B A3B — nvidia/nemotron-3-nano-30b-a3b
Llama 3.3 Nemotron Super 49B v1.5 — nvidia/llama-3.3-nemotron-super-49b-v1.5
GPT-OSS 120B — openai/gpt-oss-120b
GPT-OSS 20B — openai/gpt-oss-20b
Qwen3 Coder 480B A35B Instruct — qwen/qwen3-coder-480b-a35b-instruct
Qwen3 Next 80B A3B Instruct — qwen/qwen3-next-80b-a3b-instruct
Qwen3 Next 80B A3B Thinking — qwen/qwen3-next-80b-a3b-thinking
DeepSeek V4 Flash — deepseek-ai/deepseek-v4-flash
DeepSeek V4 Pro — deepseek-ai/deepseek-v4-pro
Step 3.5 Flash — stepfun-ai/step-3.5-flash
GLM 5.2 — z-ai/glm-5.2
GPT-5.6 Sol — gpt-5.6-sol
GPT-5.6 Terra — gpt-5.6-terra
GPT-5.6 Luna — gpt-5.6-luna
Claude Fable 5 — claude-fable-5
Claude Opus 5 — claude-opus-5
Claude Sonnet 5 — claude-sonnet-5
Claude Haiku 4.5 — claude-haiku-4-5
Gemini 3.7 Flash — gemini-3.7-flash
Gemini 3.6 Flash — gemini-3.6-flash
Gemini 3.5 Flash — gemini-3.5-flash
Gemini 3.5 Flash-Lite — gemini-3.5-flash-lite
Gemini 3.1 Pro (Preview) — gemini-3.1-pro-preview
Mistral Large (Latest) — mistral-large-latest
Mistral Small (Latest) — mistral-small-latest
Llama 3.1 8B Instruct — meta-llama/Llama-3.1-8B-Instruct
Llama 3.3 70B Instruct — meta-llama/Llama-3.3-70B-Instruct
Llama 4 Scout 17B 16E Instruct — meta-llama/Llama-4-Scout-17B-16E-Instruct
Llama 4 Maverick 17B 128E Instruct — meta-llama/Llama-4-Maverick-17B-128E-Instruct
+ Eigenes Modell hinzufügen …
Statischer Katalog, Stand 20.08.2026. Die ID wird unverändert an den Endpoint übergeben.
Endpoint-Typ
Chat Completions
Responses API
Text Completions
Embeddings
Chat Embeddings
HF TEI Rankings
Cohere Rankings
Audio Transcription
Image Generation
Raw Payload
Eigener Endpoint-Pfad
Optional, z. B. /v2/generate
Tokenizer
Leer lassen für automatische Erkennung
API-Key
Wird nie gespeichert oder ausgegeben; die Ausgabe referenziert die passende Umgebungsvariable.
Request Timeout (s)
Streaming
Aktiviert TTFT- und ITL-Messung.
Serverseitige Tokenzählung
Vertraut usage.prompt_tokens und usage.completion_tokens.
Per-Chunk Usage
Für vLLM/TRT-LLM; benötigt Streaming und Server-Tokenzählung.
Request-Optionen
Extra Inputs
JSON oder key:value, z. B. {"temperature":0.2}
HTTP Header
Einer pro Zeile, z. B. X-Tenant: demo
Zurück
Weiter
→