Why Our QA/QC Agent Failed Twice: What It Actually Takes to Build an AI Agent That Works

Inside a live Vitru AI Q&A: why our first QA/QC agent blew its token budget, what actually fixed it, and the six-part blueprint behind every agent we ship.

Key Takeaways
  • Manual QA/QC doesn’t scale because the checklist lives in people’s heads. Every reviewer’s list of issues looks a little different, revisions get chased into Friday evening, and none of it gets written down anywhere the next reviewer can use.
  • Our own first attempt at an AI QA/QC agent was a real failure, not a footnote. Feeding an entire Revit model’s elements straight into an AI’s context window found real issues and still ran up a token bill nobody could justify.
  • The fix wasn’t a bigger model — it was changing what the AI was for. Teaching the AI to write a rule check once, as a script, and letting that script run the sweep turns an expensive live judgment call into a cheap, deterministic pass.
  • An agent is six specific parts, not one clever prompt. A trigger, connections into the tools it reads and writes, a rule book, a replaceable “brain,” a defined persona and workflow, and an audit log of everything it did and everyone who approved it.
  • Rules can come from anywhere and still stay tied to one project. A firm’s own standard, a building code, or a client’s checklist all get broken into the same structured rule format, and each project keeps its own jurisdiction in memory instead of applying one global rule set everywhere.
  • A manual catch doesn’t have to stay a one-off fix. Upload the review that found it, and the agent updates its own checklist so the same issue doesn’t need catching a second time.

Most conversations about AI agents skip the part where the first version didn’t work.

That wasn’t true of the September 3, 2026 Vitru AI Q&A — open to anyone working in architecture, engineering, or construction, no pitch, just questions. The session opened with a plain account of what broke the first time a QA/QC agent got built, then walked through what actually replaced it, live, on screen.

The hour covered more ground than one post can hold, but three threads carried it: the six parts every one of these agents is actually built from, the failure that forced a real architecture change, and a live run of what happens between opening a Revit model and getting a QA/QC report back. Here’s what came out of it.

Why manual QA/QC burns through an engineer’s Friday night

Every AEC firm that reviews models by hand runs into the same wall eventually. A senior engineer checking a colleague’s Revit model against a firm’s own codes and checklists is doing something highly repetitive and something highly judgment-based at the same time, and there’s no clean way to hand off only the repetitive half. What that looks like in practice, as it was described on the call, is engineers staying in the office through a Friday afternoon and evening making sure a set of revisions is actually compliant before Monday’s submission.

A second problem shows up right behind the first one. Every engineer’s list of issues looks a little different, because the checklist mostly lives in that person’s head rather than in any shared document. Two reviewers checking the same model can flag different things, and neither is technically wrong.

Once you notice that, the fix stops looking like “add an AI chatbot to the review” and starts looking like “give every reviewer the same standard, applied the same way, every time.” That’s the actual design brief a QA/QC agent has to satisfy — not smarter judgment than an engineer already has, but one consistent check the whole firm runs against.

10–15 min
What a Revit model review that used to take two full days of engineering time takes today once the flagship QA/QC agent is checking it — not because the agent works faster than a person thinks, but because a human is no longer needed for the mechanical part of the check at all

The first QA/QC agent we built actually failed

Here’s the part vendors tend to leave out of their own case studies.

The first version tried the obvious approach: hand the AI every element in the model and let it reason over all of it in one pass. It worked, in the sense that it caught real issues. It also turned every single review into a token bill nobody could justify, because the AI was re-considering the entire model, element by element, from a blank slate, every time it ran.

The next version tried delegating the load: split the model into pieces and let separate AI calls handle each slice. That brought the context problem under control, but the whole check was still too slow to fit into anyone’s actual week.

What eventually worked meant changing what the AI was being asked to do. Instead of asking an AI model to check every element in every model live, the system now uses AI to write the rule check once, as a script — and a genuinely elaborate one, because every Revit model carries its own formatting quirks and its own mix of what’s a stored parameter versus what has to be derived from the geometry. That script is what actually runs the sweep afterward. The AI’s job is authoring the check one time; running it across hundreds of elements is cheap and deterministic from there on.

It’s not a finished system, either. Every version of it so far has been an evolution toward more efficiency, not a final answer — which is a fair description of most things built with AI right now.

What’s actually inside one of these agents

Strip away the demo and an agent built for something like QA/QC comes down to six specific parts.

  • A trigger. Something has to start it — a chat message, which makes it reactive; some other event happening elsewhere in a firm’s systems; or a schedule, like asking for a given report every Monday morning.
  • Connections into the tools it already has to touch. Revit through its own API, but also whatever else a firm already runs — AutoCAD, Rhino, SketchUp, 3ds Max, email, messaging, project repositories. The agent layer sits on top of what’s already there instead of replacing any of it.
  • A brain, and a replaceable one. The underlying AI model gets treated like infrastructure rather than a fixed choice. As long as the intelligence is strong enough for the task at hand, the system runs on a different provider without anyone noticing the swap, and cost is increasingly the lever that decides which one does the work.
  • A rule book and a persona. What a firm’s own codes, standards, and checklists actually say, plus a defined role the agent is working from — acting as a QA/QC reviewer with a specific, bounded scope rather than a generalist trying to do everything at once.
  • Guardrails. The policies layered on top of the rule book: what gets masked before anything leaves a firm’s own systems, what the agent is not allowed to touch, and protection against being talked out of its own instructions.
  • An audit trail. Every prompt, every action the agent took, every action a human approved on top of it, and the tokens it spent along the way — logged in a way a firm can look back through later to see what happened and tighten the setup.

A live walk-through: from an empty chat to a QA/QC report

  1. Setup reads the model before anything else happens. Opening the tool inside Revit brings up a chat interface on the toolbar, and it doesn’t need Revit open to log in. The first agent invoked isn’t the reviewer — it’s a setup agent that reads the project straight out of the BIM model itself: walls, doors, rooms, and the key parameters attached to each. It confirms what it found with the person running it before doing anything further, and it keeps a running memory of every version of that setup, so it doesn’t re-litigate an option that already didn’t work on an earlier pass.
  2. The agent proposes rules before it checks anything. Once the project’s shape is confirmed, the system spins up proof-of-concept rules pulled from the relevant building code — something like door marks that have to be present and unique, or fire-rated doors that need a declared width matching what the code requires. It writes the check for each rule, applies it to the file, and runs it inside the model to test it, then hands back a full report of what it proposed and tested before a single QA/QC pass has actually happened.
  3. A coordinator hands the sweep to sub-agents by discipline. With the rule set confirmed, the same session moves into the real QA/QC check. A coordinator agent doesn’t review the model itself — it spins up one sub-agent per discipline, all running in parallel, each one reporting its findings back to the coordinator. The coordinator checks each result against the rule it belongs to, marks it passed or failed, and merges everything into one final report.

Rules can come from anywhere, but they stay tied to one project

Any document can become a rule. A firm’s internal standard, a specific building code, a client’s own checklist — the agent breaks whatever it’s given into a structured rule format rather than treating it as reference text to search through later. Where a rule has a deterministic answer, the AI writes it as a script that runs inside Revit going forward, instead of forming a fresh opinion on it every time the check runs.

That doesn’t flatten every project onto one shared rule set. Each project keeps its own jurisdiction attached in the system’s memory, so the rules that actually get applied are whichever ones that project sits under. A firm working across multiple codes or regions doesn’t have to pick a single standard and apply it everywhere — the agent already knows which project it’s looking at.

Feeding a manual catch back into the agent

When a human reviewer catches something the agent missed during an ordinary manual round — not a special exercise, just a normal review — that correction doesn’t have to stay a one-off fix. Upload the review, and the agent updates its own checklist and memory from it, so the same issue doesn’t need catching a second time.

Some firms go a step further and turn their own Revit authoring policy into a gate: an engineer simply can’t submit a model until it clears the checks the firm has already defined. Combined with a checklist that grows every time a human catches something new, every next submission comes in a little cleaner than the one before it, and engineers spend less time on rework because the check is drawing on the firm’s entire accumulated knowledge base rather than one reviewer’s memory.

A GCC head start, and where the habit still breaks down

One observation from the call is worth flagging on its own. The Gulf region has one of the highest rates of Revit and BIM adoption in the industry — which is exactly why an agent that works from the model directly, rather than from drawings, has room to matter there first. Move further west and plenty of consultants and subcontractors still hand off 2D drawings out of habit, even when a Revit model already exists upstream of them. That means someone downstream ends up rebuilding a model from a drawing that was itself derived from a model in the first place.

Closing that loop — sharing IFC models directly between firms instead of flattening them into 2D and back again — is a bigger efficiency question than any single agent can answer. But it’s the direction the tooling is built to support as more of the industry works this way.

Quick tips for pointing AI at your own QA/QC process

  • Don’t hand an AI model your entire dataset and hope. If a check has a deterministic answer, have the AI write it as a script once instead of re-reasoning over the whole model on every run.
  • Split by discipline, not by convenience. A coordinator handing one lane of the rule set to each sub-agent stays both fast and scoped to what a reviewer actually needs to see.
  • Feed manual corrections back in. A review round a human catches by hand is free training data for the agent’s own checklist, if someone takes the extra step of uploading it.
  • Attach a jurisdiction to the project, not to the whole firm. The same rule engine can serve multiple codes at once if the jurisdiction lives in each project’s own memory.
  • Log everything, not just the output. Every prompt, every action taken, every human approval, and every token spent belongs in the same auditable record — not just the final report.

FAQ

What’s the actual difference between an AI agent and a smart chatbot?

A chatbot answers what you ask it and stops there. An agent perceives information from a system on its own — a Revit model, an inbox, a project board — decides what to do about it, and writes the result back into that system, whether that’s flagging an element, updating a record, or sending a message. The write-back is what makes it an agent instead of a conversation.

Why would an AI check of a single model ever cost too much to run?

It happens when an AI model is asked to hold and reason over every element of a model at once, from scratch, on every single run. That’s exactly what made our own first QA/QC attempt fail on token cost. The fix was having the AI write the rule check once, as a script, and letting that script run the actual sweep — the expensive part only has to happen a single time per rule, not once per review.

How does an AI agent handle building codes that differ by jurisdiction?

Any code, standard, or checklist gets ingested and broken into the same structured rule format, regardless of where it comes from. Each project keeps its own jurisdiction in the system’s memory, so the rule set that actually applies to a check is whichever one that specific project sits under — a firm doesn’t have to standardize on one code across every job it runs.

Can our own engineers’ manual corrections actually improve the agent over time?

Yes, if the correction gets uploaded rather than just fixed and forgotten. The agent updates its own checklist and memory from a manual review round, so the same issue is checked for automatically on the next submission instead of needing a person to catch it again.

At what point does it make more sense to run this fully in-house instead of on a commercial AI provider?

Roughly around 150 employees is where deploying a local model on a firm’s own private cloud starts making economic sense on its own terms. Below that, commercial providers already exclude enterprise account data from model training by contract — a risk profile similar to the one most firms already accept using Microsoft or Google for email.

The takeaway

What finally made a QA/QC agent work wasn’t a better model or a longer prompt. It was accepting that the AI’s best contribution is writing the check once and letting deterministic code carry it from there — then wrapping that with the parts nobody demos: a trigger, real connections into Revit and everything around it, a rule book tied to the right jurisdiction, guardrails, and a log of everything that happened. Two failed versions is what it took to learn that, and the version running today is still an evolution rather than a finished answer.

Ready to see this on your own workflow? Reserve your spot at the next weekly session — it’s open, live, and free, with no pitch. Or go straight to your own models and book a 1:1 walkthrough with our team, and we’ll map where an AI QA/QC agent actually pays off first for your firm.

Reserve your spot → Book a 1:1 →

Book a demo

See it on your model — and we’ll set up your rule package.

We’ll showcase VitruAI, run a QA pass on a model you share, learn your firm’s requirements, and stand up your first rules.

Prefer email? Send us the details and we’ll reply within one working day.

No marketing automation list. We reply by hand within one working day.