- A widely cited MIT study on enterprise generative AI found that roughly 95% of pilots never show up in the P&L. Ask most AEC firms how many of their AI experiments are still running a year later, and the pattern looks the same.
- Buying beats building, by a wide margin. A purpose-built vendor tool succeeds roughly twice as often as the equivalent built in-house — about 67% versus 33%, per that same research.
- The firms actually scaling AI aren’t doing anything exotic: they set clear risk tiers instead of case-by-case approvals, fix their underlying data before automating it, and measure return by what gets redeployed, not just what gets cut.
- The real dividing line is agentic AI versus prompting. A chatbot answers a question once, in a session, and forgets. An agent perceives, plans, acts, and remembers — across every project, not just the one it was tested on.
- We built Vitru around that distinction on purpose. The architecture behind our QA/QC agents — delegated sub-agents, deterministic rule scripts, versioned memory — is what a system running in production looks like, not a chat window bolted onto Revit.
Everyone has run a pilot. Almost nobody has a system.
By this point in 2026, most AEC firms have tried something. A business developer ran a proposal draft through ChatGPT. A BIM manager tested a Copilot plug-in on one project. Someone in leadership sat through a vendor demo, signed off on a pilot, and moved on to the next initiative. Ask how many of those experiments are still running, unattended, a year later, and most rooms go quiet.
That isn’t an AEC-specific problem. A widely cited MIT study on enterprise generative AI found that around 95% of pilots fail to produce a measurable dent in profit and loss. The researchers didn’t pin the blame on the underlying models. They pointed to something more mundane: generic tools that never learn a firm’s specific workflow, budgets aimed at the wrong department, and pilots that stayed pilots because nobody built the infrastructure to make them permanent.
AEC has its own version of the stakes attached to that first number. Deloitte’s 2026 industry outlook puts the coming labor shortfall at up to two million skilled workers by 2028. A firm doesn’t have much room to treat AI as a side experiment when the workforce math already doesn’t add up — every pilot that never scales is capacity the firm needed and didn’t get.
Why the pilot stalls
Line up the failure patterns from the research against what’s actually happening inside AEC firms, and they land in the same handful of places.
A prompt isn’t a workflow
Typing a question into a chat window and getting a good answer back feels like progress. It rarely compounds into anything. Nothing gets remembered between sessions, nothing gets reused across projects, and every new person on the team starts back at zero, re-discovering the same prompt someone else already wrote three months earlier. A pilot built entirely out of individual prompting sessions has no mechanism to become a system — it’s a skill a few people happen to have, not an asset the firm owns.
Nobody decided who owns the risk
Ask a construction executive about AI risk and you’ll often get some version of what one Nordic CEO said at the AI in AEC conference in Helsinki this year.
We know the risks of construction, but not the risks of new technology.
— A Nordic construction CEO, AI in AEC conference, Helsinki 2026
That uncertainty usually resolves one of two ways, and both stall the pilot. Either everything gets blocked by default, so the initiative quietly dies waiting on an approval that never comes, or everything gets allowed, someone eventually gets burned by a bad output touching a live project, and the whole program gets shut down in response. Neither outcome is really a technology failure. Both trace back to a decision framework that should have existed before the pilot started.
The data wasn’t ready
AI pointed at disorganized data produces disorganized output, just faster. One UK contractor, speaking at that same Helsinki conference, described spending two years organizing its safety data before that data became useful to any AI system at all. That’s not a throwaway detail — it’s the actual bottleneck. Firms that skip straight to a flashy pilot without doing that groundwork get a good demo and nothing that survives contact with a second project.
The budget went to the wrong place
The MIT research surfaced something specific worth sitting with: over half of enterprise AI spending goes toward sales and marketing tools, while the highest measured returns show up in back-office automation. AEC firms tend to mirror this. Design-generation demos get the budget and the attention because they’re visible and easy to show a client. QA/QC, proposal operations, and the CRM nobody’s touched in six months rarely get the same investment — even though that’s where the research says the money actually gets made back.
What separates the firms that scale
The firms that get past the pilot stage aren’t running fundamentally different technology than everyone else. They’ve made a handful of organizational decisions the ones stuck at the demo stage haven’t.
Risk tiers, not case-by-case
A governance model sometimes called traffic-light governance sorts initiatives into three lanes: low-risk work that proceeds without a special approval, moderate-risk work that needs a named decision-maker to sign off, and high-risk work that’s blocked outright. Leadership and IT set those checkpoints together, ahead of time.
Template libraries, not folk knowledge
A prompt or workflow that works gets written down somewhere the next project team can find it, instead of living in one person’s chat history. That’s the difference between “someone here is good at this” and “the firm knows how to do this.”
Data discipline before automation
The firms getting real output invested in cleaning and structuring their underlying data before pointing AI at it — the two-year safety-data project, not the two-week pilot. Slower up front, faster everywhere after.
Redeploy, don’t just replace
The firms that keep executive support past year one measure return by what freed-up time gets spent on next, not only by what got automated away. That’s a very different pitch to a board than a headcount line item.
What this looks like on an actual project isn’t glamorous. A Shanghai-based firm cut person-hours by roughly 80% on state grid substation work by auto-generating BIM for repetitive, standards-compliant layouts — design that follows a fixed specification rather than requiring a creative decision. A similar rule-based workflow now produces construction-ready BIM for standardized KFC outlets across China; building that script took about six months before it paid off, and it’s paying off on every site built with it since. Neither is a headline AI story. Both are systems still running well past the point most pilots quietly get abandoned.
The technical distinction underneath all of this
Everything above traces back to one distinction that gets blurred constantly: agentic AI is not the same thing as prompting a chatbot.
A generative AI tool responds to a prompt, produces an output, and waits for the next input. It doesn’t hold a goal, and it doesn’t act on anything beyond the text it returns. An agent works differently: it perceives whatever triggered it, plans a sequence of steps toward a goal, acts using real tools — updating a record, opening a file, running a check — and then evaluates whether the goal was actually met before going back to sleep until the next trigger. Chain a few of those agents together, each scoped to a narrow job, and you get what’s usually meant by “agentic AI” — not one clever tool, but an orchestrated system.
One-off prompting: answers a single question, in a single session. Nothing persists. Every user starts over, and the result depends on how good that day’s prompt happened to be.
Agentic AI in production: triggered automatically, remembers across projects, delegates work to scoped sub-agents, and gets checked before anything touches a live model or a client record.
That’s also the best explanation for why buying from a specialized vendor succeeds roughly twice as often as building in-house. A single good prompt is easy to write internally. A production agentic system — with memory that persists, delegation across sub-agents, a versioning layer, and guardrails against a bad output touching a live project — is genuinely hard infrastructure to stand up casually, on the side, without it being someone’s full-time job.
What this looks like inside an AEC firm specifically
We ran into this exact split building our own QA/QC agents. Running a full AI reasoning pass against every door, wall, and stair in a Revit file, on every revision, doesn’t hold up economically or reliably at firm scale — that’s the expensive-prompting trap in a different shape. What we built instead delegates the expensive reasoning to scoped sub-agents once, when a rule is being authored, then runs a deterministic script — not a live AI judgment call — on every check after that. A firm’s standards live in versioned memory instead of one senior reviewer’s head. That’s what a system past the pilot stage actually looks like in this industry: cheap enough to run on everything, and trusted enough that people act on what it finds.
Not sure whether your AI plan is a pilot or a system. Book a demo and we’ll walk through what’s actually running versus what’s still a demo — and what it would take to close the gap.
Book a demo → See how Vitru works
FAQ
Is agentic AI just a fancier chatbot?
No. A chatbot responds to a prompt and stops. An agent perceives a trigger, plans a sequence of actions, uses tools to act, checks whether it succeeded, and remembers the outcome for next time — without a person prompting each step.
Why do internally built AI tools fail more often than vendor tools?
Mostly infrastructure, not ambition. A production agentic system needs persistent memory, delegation across scoped sub-agents, versioning, and guardrails — real engineering work that’s hard to sustain as a side project inside a firm whose core business isn’t software.
What should a firm fix before trying to scale AI past a pilot?
Its data. Every example of a pilot that scaled involved cleaning up the underlying data first — safety records, standards, past proposals — even when that took months and produced nothing visible on its own.
Does agentic AI mean replacing our existing BIM or CRM software?
No. The agents that actually work connect to the systems a firm already runs — Revit, Airtable, a CRM — rather than replacing them. The agent acts as another user of the existing system, not a rip-and-replace.
Who should decide what AI risk a firm is willing to take?
Leadership and IT, together, before a pilot starts — not whoever happens to be running the pilot. IT usually doesn’t have the authority to grant access to sensitive systems on its own, and leadership usually doesn’t have the technical context to know what’s actually being asked for.
The takeaway
Most of what separates a scaled AI program from an abandoned pilot has nothing to do with which model a firm chose. It’s a handful of decisions made before the first agent ever touched a live project: who owns the risk call, whether the data was ready, and whether the win got written down somewhere the next project could reuse it.
Start with one workflow that already causes real pain, and build it as a system instead of a habit. Book a demo and we’ll help you figure out which one.