PUB-016/Benchmark/Aug 2026
A practical benchmark for spatial workflows
A task-based evaluation method for complete spatial workflows, covering grounding, completion time, evidence quality, privacy, accessibility, cost and human approval.
Business value
Workflow benchmarks help organisations compare systems on the constraints that affect adoption, rather than treating a model leaderboard as a deployment decision.
Evidence and measures
- - Candidate tasks include locating an asset, reviewing a hazard, routing a crew and producing a client-ready summary.
- - Scoring dimensions include spatial accuracy, task completion, latency, evidence quality, explanation and failure recovery.
- - The benchmark records privacy mode, deployment path, accessibility and the point where human approval is required.
Research Questions
- 01Can a small shared task set make spatial workflows comparable across different sectors?
- 02How should the score change when a workflow observes, recommends, drafts or acts?
- 03Which measures best predict whether an operator will trust and reuse the system?
Method
- 01Start with a defined operator job rather than an abstract model prompt.
- 02Record data, device, model, tools, timing conditions and human-review conditions for each run.
- 03Score the complete path from spatial evidence to approved handoff, including failure behaviour.
Evidence
- 01Simam research note: A practical benchmark for spatial workflows.
- 02Existing Connected Highways, CivilMap and venue-management task patterns.
- 03Simam benchmark method: best fit, operational speed, privacy, evidence quality and blueprint fit.
Next Steps
- Define a public-safe four-task benchmark pack.
- Run the same tasks across map-first, 3D-first and report-first interfaces.
- Publish failure cases and measurement conditions alongside any score.
Related live work
Related Benchmark records