Private beta Now onboarding founding customers — preferred rates and first say on the roadmap. Join the beta →
  1. Overview
  2. Benchmark

ConstructBench — our benchmark for construction AI

An AEC benchmark focused on construction and real-world applications: real job documents and drawings, scored the way a contractor would judge the answer. Coming soon, with our own models' results against the leading general-purpose models published as each run lands.

Coming soon

An AEC benchmark focused on construction and real-world applications — every task drawn from the documents, drawings and decisions a contractor actually works with, and scored the way a contractor would judge the answer, so a model that is genuinely good at construction shows it and one that only sounds confident does not.

Leaderboard

Our models against the leading general-purpose models. Each cell publishes the day its run lands.

Construction benchesSix groups, drawn from real job material
Results publish as runs land
BenchWhat it scoresConstructelligenceFrontier general-purpose
Estimating & takeoffQuantities read from a drawing set and priced against a cost codeComing soonComing soon
SchedulingActivity durations and sequence from a scope, a crew and the site conditionsComing soonComing soon
Cost & forecastingCost at completion from a job's own burn — the number the platform turns onComing soonComing soon
DocumentsNotice periods, exclusions and the clauses that do not flow down, read from real contractsComing soonComing soon
Field & progressWhat is actually installed, read from logs, photos and quantitiesComing soonComing soon
Reference dataCost codes, units and waste factors — the data every other answer depends onComing soonComing soon

Every run publishes with the harness, the prompts and the scoring code beside it, so anyone can re-run it and get the same number.

How it is scored

Real job material

The sample set a contractor would recognise — plans, estimates, contracts, schedules and job cost — not puzzles written for a model.

A contractor's tolerance

A quantity or a cost is right, within a stated band, or wrong. Partial credit is explicit, never hidden inside an average.

Per task and overall

A model that is strong at takeoff and weak at forecasting is visible as exactly that, not smoothed into one score.

Why we publish it

The forecasting models behind the product are trained on proprietary construction cost data and stay private. ConstructBench is the public side of that work: an open harness, openly scored, so a buyer can tell a real construction model from a general-purpose one wearing a hard hat.

What stays open, and what does not

The tasks, the harness, the scoring code and the results are open. The dataset behind our own forecasting models is not: it is licensed and private, and ConstructBench is built on a separate, publishable sample set made with the same method.

Want early access to the results? Join the private beta Follow on GitHub Follow on Hugging Face

More of the platform

Your numbers are already there. Go look at them properly.

Hosted for you or self-hosted, read-only, every user included. Connected and reconciled against your own reports in a single 90-minute session.

Private preview

Request access to the demo

The full dashboard on the eight-job demo portfolio is invite-only for now. Leave your work email and we will send you access within one business day.

Already have the password? Open the demo →

Join the private beta Demo access