Critical
Urgent
Low
Resolved
📋Run a scenario to see the incident timeline.
What's being measured?
Quality Score (0–100) — decision correctness: rule compliance, resolution rate, escalation accuracy. This is what matters for judges.
Avg Response Time — API wall-clock time. Multi-agent = 4 Qwen calls per incident vs 1 for single-agent. It's expected to be slower. In real deployment, agents would run as parallel microservices.
Priority Violations — did a LOW incident get a resource while a CRITICAL one was waiting? The Auditor prevents this. The single agent has no one double-checking.
📊Run "Full Comparison" to see the benchmark.