Field research
Agentic AI for the built environment
Quantified models of where agents actually change the economics of development, design and construction across Ireland and the UK. Every figure states its assumptions.
Real Estate & Development
10 min read
A number is not an estimate: the multi-agent study that produced cost distributions instead
Twelve cost engineers, three real buildings, and a measured comparison against deterministic BIM, manual probabilistic estimating and single-prompt GPT-4. The agents won, and the ablation says why.
12.5%
MAPE, against 18.7 to 27.4%
4.2 min
Per estimate, from 18.5 to 68
$0.15
API cost per run
9 min read
Site selection in eleven days: running development due diligence with agents
A worked model of agentic site screening for Irish and UK development, including where the time actually goes and which parts refuse to automate.
14 wks → 11 days
Brief to ranked shortlist
63
Sites screened per cycle
€31k
Modelled saving per scheme
8 min read
The lease portfolio that reads itself: compliance agents across commercial estates
Rent review triggers, break clauses and statutory deadlines are a data extraction problem that portfolios keep solving with calendar reminders and hope.
1,240
Leases parsed per run
94%
Clause extraction accuracy
£410k
Missed escalation exposure found
Architecture & Predictive Design
10 min read
Let the agent write code: TUM's result that the paradigm beats the model
A 1,027-task benchmark across 37 IFC models found that letting an agent explore a model by running code beat static query generation by 37 points. The weaker adaptive model beat the stronger static one.
+36.8 to 38.5pp
Adaptive over static, p < 0.001
1,027
Benchmark tasks, 37 models
55 to 57%
Strict accuracy, still not enough
8 min read
A 7B model that beat Claude and GPT-5.2 at reading building regulations
Fine-tuning plus reinforcement learning on a narrow regulatory task outperformed frontier models zero-shot. The argument for training rather than prompting, and for keeping regulation on your own hardware.
−23.8%
Tree edit distance vs SFT baseline
−38.6%
Token-level Levenshtein distance
7B
Parameters, locally hostable
7 min read
Five papers, one pattern: nobody won by using a bigger model
Reading a month of agentic AI research in the built environment side by side. The architectures differ completely and the conclusion is the same one every time.
5 papers
Read side by side
0
Won by scaling the model
2
Where data quality set the ceiling
10 min read
Every facade option, priced four ways: predictive design past generative design
Generative design produced options nobody could evaluate. Predictive design forecasts how each option performs structurally, financially and in embodied carbon before anyone commits.
4 min
Per option, fully appraised
217 kgCO₂e/m²
Spread across options
6.2%
Capital cost variance found
9 min read
Agents that drive the CAD software: what Model Context Protocol changes for design practices
The interesting shift is not agents that describe engineering work. It is agents that operate the specialist tools directly, and the governance that has to come with it.
71%
Modelling tasks agent-operable
3.4×
Option throughput at stage 3
100%
Actions requiring audit trail
Construction Management
8 min read
One photo, one updated model: what the first MCP progress-monitoring paper actually shows
A research team wired Claude to Revit and Blender through MCP and closed the loop from site photograph to updated 4D model with nobody in the middle. The result is an existence proof, not a product.
4 stages
Fully autonomous loop
91.4%
Completion computed from one image
1 wall
Total experimental scope
9 min read
The front half of compliance checking is nearly solved. The back half is where the work is
97% F1 on rule classification and 100% on dependency identification. Interpreting regulations turns out to be the easy part, and the industry has been optimising the wrong end.
97%
F1, rule classification
100%
Dependency identification
97%
Correct tool selection
8 min read
The first draft of every RFI: document agents on live construction projects
Site teams lose more hours to structured paperwork than to any technical problem. That work has a shape, and the shape is automatable.
11.4 hrs
Weekly admin recovered per PM
2.1 days
RFI turnaround, from 6.8
412 pp
Tender reviewed in one pass
9 min read
Closing the loop between site and model: reality capture as a live data asset
Drone and 360 capture produces enormous datasets that mostly get archived. Agents turn that data into deviation detection, and deviation detection into design action.
8 days → 4 hrs
Capture to deviation report
23mm
Detectable deviation threshold
£1.9m
Modelled rework avoided
Want this modelled against your own numbers?
Every model on this page is built from stated assumptions and takes an afternoon to re-run on a real portfolio, pipeline or project.