<span class="vcard">/u/MuhammadMujtaba21</span>
/u/MuhammadMujtaba21

I’m building an independent verification layer for Ai generated-claims and I’m lokking for researchers and partners to build with us.

I've been working on a deterministic verification engine for AI-generated financial claims. The original idea was fairly simple: An LLM should generate claims. It shouldn't be the authority that verifies them. But after building and testing the…

I reran the benchmark. The deterministic result reproduced exactly — but the model-related metric tells a different story.

I reran the benchmark. The deterministic result reproduced exactly — but the model-related metric tells a different story. After the discussion on my previous benchmark, I reran the verification capability benchmark and inspected the results more caref…

Our deterministic verification engine passed 66/66 benchmark cases on canonical structured inputs.

Our deterministic verification engine passed 66/66 benchmark cases on canonical structured inputs. In live model evaluation, the end-to-end pipeline currently passed 19/66 cases. We are restructuring the benchmark to isolate failures by their first inv…

I benchmarked my deterministic AI financial verification engine. The core passed 66/66, but the live LLM pipeline only passed 19/66.

I've been building a deterministic verification engine for AI-generated financial claims. The basic idea is simple: An LLM can generate a financial answer, but the LLM itself should not be allowed to decide that its answer is "verified.&…

Update: posted here asking what would make you trust AI financial calculations. The best critique broke my core assumption — here’s what changed.

A little while back I posted here asking accountants what it would actually take to trust an AI-generated financial calculation. I said I was looking for reasons not to pursue this, not encouragement. You delivered — genuinely the sharpest feedback I&#…

I built a deterministic engine that catches AI’s financial math errors before they ship — looking for people to poke holes in it

Quick context: I've spent the last year+ building something in the "AI hallucination" space, specifically for finance, and I want honest feedback before I go further — not upvotes, actual criticism. The problem I'm trying to solve: AI…

How would answer these?

What is a claim? What is evidence? What is a constraint? What is a proof? What is an assumption? What is a contradiction? What is trust? submitted by /u/MuhammadMujtaba21 [link] [comments]

What does it mathematically mean for an AI-generated claim to be "true", "justified", and "trustworthy"?

I'm working on a research project, the end goal of which is not to create a better LLM, but rather to create a verification engine that can reason about whether an AI claim is trustworthy enough for a particular application. Most of the current res…

We’re trying to answer a simple question: Can AI prove it’s right before you trust it

​ Over the last few months, I've been building AutoFlow, not as another AI wrapper or workflow tool, but as a verification engine.Instead of asking: "What does the model think?" we're asking: Can the answer be mathematically, l…

Title I’m looking for engineers who enjoy solving problems that are more about correctness than AI.

Over the last few months I've been building a prototype around a question I can't stop thinking about: How do you know when an AI-generated financial claim is actually trustworthy? The obvious answer is "use a better model." The more …