back to portfolio

Field research

Agentic AI for the built environment

Quantified models of where agents actually change the economics of development, design and construction across Ireland and the UK. Every figure states its assumptions.

Architecture & Predictive Design

10 min read

Let the agent write code: TUM's result that the paradigm beats the model

A 1,027-task benchmark across 37 IFC models found that letting an agent explore a model by running code beat static query generation by 37 points. The weaker adaptive model beat the stronger static one.

+36.8 to 38.5pp

Adaptive over static, p < 0.001

1,027

Benchmark tasks, 37 models

55 to 57%

Strict accuracy, still not enough

Read

8 min read

A 7B model that beat Claude and GPT-5.2 at reading building regulations

Fine-tuning plus reinforcement learning on a narrow regulatory task outperformed frontier models zero-shot. The argument for training rather than prompting, and for keeping regulation on your own hardware.

−23.8%

Tree edit distance vs SFT baseline

−38.6%

Token-level Levenshtein distance

7B

Parameters, locally hostable

Read

7 min read

Five papers, one pattern: nobody won by using a bigger model

Reading a month of agentic AI research in the built environment side by side. The architectures differ completely and the conclusion is the same one every time.

5 papers

Read side by side

0

Won by scaling the model

2

Where data quality set the ceiling

Read

10 min read

Every facade option, priced four ways: predictive design past generative design

Generative design produced options nobody could evaluate. Predictive design forecasts how each option performs structurally, financially and in embodied carbon before anyone commits.

4 min

Per option, fully appraised

217 kgCO₂e/m²

Spread across options

6.2%

Capital cost variance found

Read

9 min read

Agents that drive the CAD software: what Model Context Protocol changes for design practices

The interesting shift is not agents that describe engineering work. It is agents that operate the specialist tools directly, and the governance that has to come with it.

71%

Modelling tasks agent-operable

3.4×

Option throughput at stage 3

100%

Actions requiring audit trail

Read

Want this modelled against your own numbers?

Every model on this page is built from stated assumptions and takes an afternoon to re-run on a real portfolio, pipeline or project.