What does it mathematically mean for an AI-generated claim to be "true", "justified", and "trustworthy"?
What does it mathematically mean for an AI-generated claim to be "true", "justified", and "trustworthy"?

What does it mathematically mean for an AI-generated claim to be "true", "justified", and "trustworthy"?

I'm working on a research project, the end goal of which is not to create a better LLM, but rather to create a verification engine that can reason about whether an AI claim is trustworthy enough for a particular application.

Most of the current research focuses on making AI "better", I want to tackle the verification side of things.

The question I'm asking myself is:

What does it mean for a claim to be "true", "justified", and "trustworthy"?

I'm not satisfied with the philosophical answers, I want to see mathematical formalisms.

Some of the questions I'm trying to answer are:

Can "trust" be formalized as a function? How is it related to truth, evidence, proof, constraints, uncertainty?

Should it be approached from the angles of probability theory, information theory, formal logic, graph theory, topology, category theory, optimization, etc.?

Can one represent any claim as an object with evidence, assumptions, constraints, and derivations?

Is there existing work on proving claims of AI (not just trusting the model's "confidence")?

How would you differentiate between a true claim, a justified claim, and a trustworthy claim from a mathematical point of view?

How would you design a Trust Engine if you had to build it from scratch? What mathematical foundations would you use?

What I'm thinking about is something akin to constraint satisfaction, where a claim needs to satisfy all constraints (logical, mathematical, evidential) to be considered trustworthy. Another approach is to think of trust as a limiting case of evidence, but I'm not sure if that's a mathematically sound way to reason about it.

I'm asking for recommendations on papers, books, etc., related to the topics.

I'm also asking for potential pitfalls in my thinking. What's wrong with the ideas I've stated above?

I'm most interested in responses from people working in formal methods, theorem proving, mathematical logic, knowledge representation, verification, optimization, information theory, and trustworthy AI.

I'm especially interested in hearing how you would approach the Trust Engine design from first principles.

submitted by /u/MuhammadMujtaba21
[link] [comments]