The sample set a contractor would recognise — plans, estimates, contracts, schedules and job cost — not puzzles written for a model.
- Overview
- Benchmark
ConstructBench — our benchmark for construction AI
An AEC benchmark focused on construction and real-world applications: real job documents and drawings, scored the way a contractor would judge the answer. Coming soon, with our own models' results against the leading general-purpose models published as each run lands.
An AEC benchmark focused on construction and real-world applications — every task drawn from the documents, drawings and decisions a contractor actually works with, and scored the way a contractor would judge the answer, so a model that is genuinely good at construction shows it and one that only sounds confident does not.
Leaderboard
Our models against the leading general-purpose models. Each cell publishes the day its run lands.
| Bench | What it scores | Constructelligence | Frontier general-purpose |
|---|---|---|---|
| Estimating & takeoff | Quantities read from a drawing set and priced against a cost code | Coming soon | Coming soon |
| Scheduling | Activity durations and sequence from a scope, a crew and the site conditions | Coming soon | Coming soon |
| Cost & forecasting | Cost at completion from a job's own burn — the number the platform turns on | Coming soon | Coming soon |
| Documents | Notice periods, exclusions and the clauses that do not flow down, read from real contracts | Coming soon | Coming soon |
| Field & progress | What is actually installed, read from logs, photos and quantities | Coming soon | Coming soon |
| Reference data | Cost codes, units and waste factors — the data every other answer depends on | Coming soon | Coming soon |
Every run publishes with the harness, the prompts and the scoring code beside it, so anyone can re-run it and get the same number.
How it is scored
A quantity or a cost is right, within a stated band, or wrong. Partial credit is explicit, never hidden inside an average.
A model that is strong at takeoff and weak at forecasting is visible as exactly that, not smoothed into one score.
Why we publish it
The forecasting models behind the product are trained on proprietary construction cost data and stay private. ConstructBench is the public side of that work: an open harness, openly scored, so a buyer can tell a real construction model from a general-purpose one wearing a hard hat.
What stays open, and what does not
The tasks, the harness, the scoring code and the results are open. The dataset behind our own forecasting models is not: it is licensed and private, and ConstructBench is built on a separate, publishable sample set made with the same method.
Want early access to the results? Join the private beta Follow on GitHub Follow on Hugging Face
More of the platform
Your numbers are already there. Go look at them properly.
Hosted for you or self-hosted, read-only, every user included. Connected and reconciled against your own reports in a single 90-minute session.