August 2026
Can an AI Agent Run Your Month-end Accounting Close?
An honest account of a pilot that didn't work, and what it proved
Most AI pilot updates are progress reports. This one’s an honest account of a pilot that didn't work, told by the two people closest to it: the tester who ran it, and the leader now responsible for acting on what it revealed.
Earlier this summer, Mabela Zenullari, who leads client implementation quality across several of Propeller's industry verticals, tested whether Claude could run a real client's monthly close. Zain Khoja, Propeller's VP of Research & Development, owns what happens next. We caught up with the two of them after they presented the results at a company all-hands meeting.
How do you test whether an AI agent could run a real client's monthly close? Walk us through the setup.
Mabela: The setup wasn’t to prove that the agents could work; it was to look at whether they would fail. We took a real client month-end reconciliation and reconstructed the books the way they looked before the close: every bill entered, all the operational activity, nothing cleaned up in advance. As the answer key for this test, we used the close our controllers had actually booked. Then, we let the agent run on the pre-close data and graded what it produced against what we really booked. No vibes, no demo magic; just the output against the actual truth.
What did you find?
Mabela: What came back didn't surprise me or the other testers, and it wasn't “bad”; the problem was that it didn’t actually save us any time. Checking the agent's work took about as long as doing the close manually. The agents were improvising each reconciliation, finding their own way through: usually a reasonable way, but their own. So every result had to be validated from scratch against a person's judgment, not against a standard. That's not automation: it’s just a second set of books to review. Editor’s note: Workday found a milder version of the same thing across 3,200 workers: for every ten hours of efficiency gained through AI, nearly four are lost to fixing its output. In our pilot example, the clawback was closer to all ten.
If the output wasn’t wrong, why do you consider this AI close pilot a failure?
Mabela: When the output was off, we couldn't tell whether the agent had failed or whether it had correctly executed a workflow that had simply never been written down. This is why workflow documentation before automation is so critical: without a written standard, every error is ambiguous. You can't debug the tool or the process, because you can't tell them apart. It's the same thing Chris wrote about a few weeks ago: automating a process requires the process to exist as a process, not as a set of habits living in someone's head. We weren't just building an answer key for this one pilot. We were building it for everything that comes after.
Is a failed AI pilot still a win?
Mabela: I'd call it a pilot that told the truth. Our first approach didn't beat manual, and we're saying that out loud on purpose. We hope it’s exactly why the second approach will earn people's trust. Documentation isn't a detour from automation; it's the part our test proved we can't skip. Once a reconciliation exists as a written standard, an agent can be checked against that standard in minutes, instead of against someone's head in hours.
Zain, this is where you pick up. How do you turn a failed AI pilot into a better plan?
Zain: Based on everything we’ve learned, the takeaway is to standardize the workflow first, then automate. So our embedded teams are codifying the month-end close process itself. Each controller has been assigned a set of GL accounts, and their job is to document the exact workflow end-to-end for each one: what inputs are needed and what steps are taken, in what order, to drive that reconciliation from start to finish. Editor's note: BCG has granular recommendations here: successful AI adoptions apply 70 percent of effort to people, process, and culture; 20 percent to data and technology; and 10 percent to the algorithms. Our pilot just ran this math in miniature.
I want to be clear about what this isn't: it's not about redesigning the reconciliation template. The goal is a workflow that's not only accurate but repeatable, because this consistency is what sets the foundation for the next step, which is automation.
Does 'document first, automate second' really work, or is it a bet?
Zain: It's not a bet. Two other workstreams in our close-automation program already prove it: an Amazon merchant-reconciliation tool and a payroll reconciliation tool, both live today, both built in exactly this order: defined workflow first, automation layered on top for the exceptions and the judgment calls. The Claude pilot is the same lesson, learned the hard way, on a process nobody had gotten around to writing down yet.
Why put a pilot that didn't work in front of the whole company, instead of quietly retiring it?
Mabela: Because I think people learn more from what didn't work than from what did. A success story tells you that someone got there; it doesn't tell you where the holes are. This pilot showed us exactly where the holes are, and that's worth more to the next team than a polished win would be. If we shelve it quietly, what we learned dies with the pilot, and the next team makes the same mistake in private instead of starting from what we already know.
Zain: And it's consistent with how we want to do this everywhere, not just in the close. An eighteen-year-old firm with hundreds of clients doesn't get automation right by pretending it worked the first time out. We get there by being precise about why it didn't work, out in the open, and turning that precision into the next thing we build.
Some AI research numbers, for context:
- “Change management due to technology and AI” is the top issue for accounting firms, according to a recent AICPA survey
- An S&P Global report found that 42% of companies abandon their AI initiatives before they go into production.
- Sixty-one percent of CFOs cite data quality as the biggest barrier to implementing AI in finance, according to an E&Y report; unclear benefits are the second greatest challenge (51% of respondents), followed by a lack of skills, resources, or capacity (50% of respondents)
- Gartner projects that fixing context around their data can improve companies' agentic AI accuracy by up to 80%; Gartner analyst Rita Sallam: “Without context — a clear understanding of the specific relationships and rules within an organization’s data — AI agents cannot operate accurately"