What is an AI twin?
The term is used loosely enough to be meaningless. Here is a definition precise enough to argue with.
An AI twin is a representation of one specific person that can answer questions on their behalf, grounded only in information that person has provided, within boundaries that person has set.
Three clauses in that sentence are doing all the work: one specific person, grounded only in, and boundaries that person has set. Drop any of them and you have something else.
What an AI twin is not
It is not a chatbot with your name on it
A general-purpose model prompted to 'act like Ali' is doing impersonation, not representation. It will answer questions about your company by producing text that sounds like the kind of thing a founder says, because that is what it was trained to do. When it does not know your revenue, it will produce a plausible number. That is not a small flaw; it is the whole difference between a party trick and something you would put your name on.
It is not a search index over your files
Retrieval alone gives you a document finder. A twin has to synthesise: connect a metric in a spreadsheet to a claim in a deck to a caveat in a memo, and produce the answer a human would have given. Retrieval is the floor, not the product.
It is not a digital clone of your personality
Nothing here reconstructs consciousness, judgement or taste. A twin handles the part of you that is reference material - what you know, what you have built, what you are looking for. The part of you that decides is not delegated and should not be.
The three properties that matter
Grounding
Every answer traces back to a document you supplied. Not 'mostly', not 'usually' - if the graph does not contain the answer, the correct output is 'I do not have that, let me pass it to Ali'. A twin that guesses is worse than no twin, because it damages the trust of exactly the people you were trying to impress.
Permissioning
Different people get different answers, deliberately. A public visitor, a contact from a conference, a named investor and a client each sit in a different zone with a different slice of your knowledge visible. This is not a privacy feature bolted on afterwards; it is the thing that makes it safe to upload anything real in the first place.
Interruptibility
You can take over any conversation at any moment. The twin is a first responder, not a replacement, and the handover has to be instant and obvious to the other party. A system you cannot interrupt is one you will never fully trust with anything that matters.
How one actually gets built
- You upload what you know: deck, memo, metrics, CV, project docs, notes.
- The engine extracts entities and relationships into a private knowledge graph - not just chunks of text, but the connections between them.
- Answers are synthesised from that graph, with a citation attached to each claim.
- A policy layer decides, per asker, which parts of the graph are in scope.
- Anything out of scope routes to you instead of being answered.
The knowledge graph is the part that distinguishes a twin from a document chatbot. A graph knows that the ARR figure in the September memo supersedes the one in the June deck, and that the customer named in both is the same entity. That structural understanding is why the answers hold up when someone digs.
What to expect, honestly
- It will be excellent at questions whose answers exist in your documents.
- It will be useless at questions about things you never wrote down.
- It will not have your judgement, and should not pretend to.
- Its quality is bounded almost entirely by what you upload.
That last point is the one people underestimate. A twin fed ten real documents outperforms one fed a hundred pages of marketing copy, because grounding rewards substance and punishes padding. The work of building a good twin is mostly the work of writing down what you actually know.