Pre-fill Synthetic Test:
Temp: 0.1
Max Tokens: 1024
Local Qwen: qwen3-8b-local (/no_think)
Executing Concurrent Inferences...
Dispatching prompt to External Provider (if configured) and Local Qwen (CPU 6-threads with /no_think)...
External Model
TTFT
-
Total Latency
-
Tokens (In / Out)
-
Tokens / Sec
-
JSON Validity
-
Est. Cost
-
Model Confidence: Not provided
Truncated: No
qwen3-8b-local (Hetzner CPU)
TTFT
-
Total Latency
-
Tokens (In / Out)
-
Tokens / Sec
-
JSON Validity
-
Est. Cost
$0.000000 (Local)
Model Confidence: Not provided
Truncated: No
Manual Reviewer Assessment
Auto-saved to historyComparison History
All historical test records stored persistently in SQLite
| Time | Workflow | External Model | Latency (Ext vs Qwen) | Tokens (Ext vs Qwen) | Ext Cost | Decision | Score | Action |
|---|---|---|---|---|---|---|---|---|
| Loading history records... | ||||||||
Aggregated Benchmark KPIs
Real-time metrics computed across all historical comparison runs
Total Comparisons
0
Persistent database records
Cumulative Ext Cost
$0.00
Qwen Cost: $0.00 (Self-Hosted)
Avg Latency (Ext / Qwen)
0.0s / 0.0s
Median: 0.0s / 0.0s
Avg TTFT (Ext / Qwen)
0.0s / 0.0s
Speed: 0 tok/sReviewer Decisions & Parity Distribution
External Better
0
Qwen Better
0
Equivalent / Parity
0
Both Unacceptable
0
Data Export & Backup
Download complete historical benchmark evaluations for offline reporting or training readiness analysis.
Server Provider Status
Local LLM: qwen3-8b-local
Qwen Network: orviq-llm (isolated bridge)
Database Storage: /app/data/compare.db (SQLite)
External Provider API Key: Checking...