Update: posted here asking what would make you trust AI financial calculations. The best critique broke my core assumption — here’s what changed.
Update: posted here asking what would make you trust AI financial calculations. The best critique broke my core assumption — here’s what changed.

Update: posted here asking what would make you trust AI financial calculations. The best critique broke my core assumption — here’s what changed.

A little while back I posted here asking accountants what it would actually take to trust an AI-generated financial calculation. I said I was looking for reasons not to pursue this, not encouragement. You delivered — genuinely the sharpest feedback I've gotten anywhere on this, and I want to close the loop on what it changed.

The critique that mattered most (paraphrasing u/usually_guilty99):

That's a direct hit on the core premise, not an edge case. A few other people independently converged on the same wall from different angles — derived figures with no clean source ("no receipts available"), as-reported vs. revised financials, and the basic point that accountants don't verify a number by recreating the whole report, they ask for workings and interrogate judgment calls.

What I got wrong in the original pitch:

I was implicitly promising "deterministic verification" as if it applied uniformly to any financial calculation. It doesn't, and pretending otherwise is worse than the problem I'm trying to solve — a confidently wrong deterministic engine is more dangerous than a confidently wrong AI, because it comes wrapped in false certainty.

What changed:

The tool now has to do something it didn't do before: explicitly say "cannot verify — no unambiguous rule/source mapping" instead of forcing a number whenever the calculation requires interpretation, judgment, or a source field that isn't cleanly defined. Determinism only gets claimed where it's actually earned. Everything else surfaces as "needs human judgment," not a confident wrong answer.

This is a real design constraint now, not a caveat in a pitch deck — it changes what the tool is allowed to output, not just how it's described.

Where it stands:

  • Deterministic verification still works end-to-end for the class of calculations where source-to-formula mapping is genuinely unambiguous (started with net leverage and a few adjacent ratios)
  • New: explicit "unverifiable" output state for anything outside that — not a forced answer, not silence, a distinct third category
  • Still open, and still the thing I'm least sure about: where exactly that line sits in practice, across different calculation types

What I still want to know, now more specifically:

  1. If you've got a real (sanitized/hypothetical is fine) example of a calculation that looks mechanical but actually needs judgment — I'd genuinely like to see it. Trying to map the actual boundary, not the one I assumed going in.
  2. For the people who said "I like the separation between AI and deterministic logic" — does that trust survive once the tool also has to say "I don't know" sometimes? Or does an "I don't know" from a verification tool undermine confidence in the cases where it does give an answer?
  3. If anyone from the original thread (or anyone new) wants to actually try breaking this on a real scenario — genuinely open to that, no pitch, no cost, I'd rather find the failure case with someone who knows what they're doing than guess at it alone.

Thanks to everyone who commented on the original post — this is a materially different (and more honest) design than what I posted a few weeks ago, and that's because of the pushback, not in spite of it.

https://www.reddit.com/r/artificial/comments/1vkqiik/comment/p2ywzdz/?screen_view_count=2

submitted by /u/MuhammadMujtaba21
[link] [comments]