AI AI basics: what it is and how it got here
What AGI means, and why smart people disagree about it
Artificial general intelligence explained without hype: the competing definitions, what today's systems still cannot do, and how to read a confident prediction about it.
The short answer
- AGI means a system that handles the full range of cognitive work a capable adult can handle, including tasks it was never trained for, and there is no agreed test for it.
- Most public disagreement about AGI is definitional: an economic definition, a task breadth definition, a learning efficiency definition and an autonomy definition can be met in different orders.
- Benchmarks cannot settle the question because published tests leak into the training data of later models, so a high score may show exposure rather than ability.
- Current systems reliably fall short on four things: learning after release, holding a consistent model of the world, staying on a long goal, and learning from few examples.
- Timelines vary because forecasters disagree about whether scaling continues, whether the missing pieces are engineering or unknown ideas, and what counts as arrival.
AGI, artificial general intelligence, means a system that can handle the full range of cognitive work a capable adult can handle, including tasks nobody trained it for. That is the shared core of every definition. Beyond that core the agreement stops, and the disagreement is not a detail. There is no accepted test, no agreed list of abilities and no measurement that would let two experts look at the same system and settle the question. When you see a confident claim that AGI is five years away or fifty, most of the distance between those numbers comes from the speaker using a different definition, not from different evidence.
The definitions people are actually using
Four definitions do most of the work in public argument, and they can be satisfied in different orders. A system could pass one and fail the rest.
| Definition | What it says AGI is | What would count as evidence | Where it gets awkward |
|---|---|---|---|
| Economic | A system that can do most of the work people are paid to do at a desk | Sustained replacement across many occupations, not demonstrations | Confuses capability with deployment, which is slowed by cost, liability and trust |
| Task breadth | Matches competent adults across the whole span of cognitive tasks | A wide battery of tasks, including ones written after the model was trained | Nobody agrees which tasks belong on the list, and the list keeps being revised |
| Learning efficiency | Picks up a genuinely new skill from about as little experience as a person needs | Teaching it an invented game or tool from a few examples, and having that stick | Hard to test, because the skill may already sit somewhere in the training data |
| Autonomy | Carries a long, messy goal to completion and recovers from its own mistakes | Multi week projects finished without a person quietly unsticking it | Results depend as much on the scaffolding around the model as on the model |
| Understanding | Genuinely understands rather than imitating understanding | No agreed test exists | Not measurable, so it can never settle an argument either way |
Notice how much rides on which one somebody picked. On the economic definition, a system that writes excellent code but cannot be trusted alone with a bank login is not AGI. On the task breadth definition it might already be close. The word carries whichever meaning the speaker needed.
Why there is no exam that settles it
The obvious fix is a test. It has been tried, and the pattern of failure is consistent.
Turing proposed the imitation game in 1950: if a person conversing by text cannot tell machine from human, call the machine intelligent. Modern systems do well at that and almost nobody accepts the conclusion, because it turned out to measure fluency and social mimicry rather than general capability. The test was retired by its own success.
Benchmarks have the opposite problem. A benchmark is a fixed set of questions, so once it is published it drifts into the training data of the next generation of models. A high score then means the answers were seen, not that the ability is there, which is why every serious evaluation now worries about contamination and why new benchmarks are built to be unguessable and get saturated anyway.
Professional exams mislead in a third way. Passing a bar exam and practicing law are different jobs. Exams are built to separate prepared humans from unprepared ones, and they assume everything an exam cannot check, such as judgment, client contact and knowing when to stop, is already in place. A machine score on a human exam does not carry those assumptions.
What current systems still cannot do
Whatever definition you favor, four gaps show up repeatedly, and none of them is a small bug.
Continual learning. A released model is frozen. Its weights stop changing when training finishes, so it cannot absorb something new from working with you. What looks like learning inside a conversation is your earlier messages still being visible, and it disappears when the chat ends or gets too long. Being unable to update is also why every model has a date after which it knows nothing.
A stable model of the world. These systems hold facts that contradict each other and do not notice, because nothing requires consistency. Ask about physical arrangements, what is inside a closed container, how objects stack, what a person in the room can and cannot see, and you get plausible answers that fall apart under a follow up. Reasoning modes that work through intermediate steps improve this measurably and do not solve it.
Goal persistence. Give a system a task with twenty steps and errors compound. A small wrong turn at step three quietly poisons everything after it, because there is no separate part of the system checking whether the plan still makes sense. This is the central practical problem in systems that take actions rather than answer questions.
Sample efficiency. A child learns a new word from a handful of exposures. A frontier model reads trillions of words to reach its level of fluency and still needs many examples to pick up an unfamiliar pattern reliably. Whether that gap is an engineering problem or a sign the approach is different in kind is one of the genuine open questions.
There is an older observation behind all of this. The things people find hard, chess and calculus and legal drafting, turned out relatively easy for machines. The things any toddler does, walking into an unfamiliar kitchen and making a sandwich, turned out very hard. Progress on the first list tells you less about the second than intuition suggests.
Why the timelines disagree so much
Serious researchers give estimates ranging from a few years to never, and when they are surveyed the spread is wide and the median date has shifted substantially between surveys. Three disagreements produce most of that range.
The first is about scaling. One camp reads the last decade as a straight line: more data, more compute and more parameters kept producing more capability, so extrapolate. The other camp points out that returns on sheer parameter count have flattened before, that high quality training text is finite, and that a curve is not a law. Both are reading the same chart.
The second is about what is missing. If continual learning and reliable world models are engineering problems, they get solved on an engineering schedule. If they need an idea nobody has had yet, no schedule applies, because you cannot put a date on an unknown insight.
The third is about what counts as arrival. Someone using the task breadth definition can reasonably name a much earlier date than someone who requires sustained economic replacement. They may not actually disagree about any fact.
Incentives sit on top of all three. A near date attracts investment and attention; a far date signals rigor and caution. Neither motive makes a prediction wrong, and both are worth noticing.
How to read a confident prediction
Five questions will tell you more than the headline number.
- Which definition is being used? If the claim does not say, it cannot be checked and probably cannot be wrong either.
- What evidence is offered, and is it a benchmark score or behavior in the world? Scores on published tests are the weakest evidence available, for the contamination reason above.
- What would change the speaker's mind? A forecast with no failure condition is a position, not a prediction.
- Is the claim about capability or about deployment? A capability can exist for years before anything in your life changes, because integration, regulation and liability move at their own speed.
- Has the speaker been specific before, and did it happen? Track records are rare in this field and worth a lot when they exist.
The same questions work in reverse on claims that nothing is happening. Dismissals are predictions too, and they are just as often built on a definition chosen to make the answer come out right.
Living with an unsettled question
You do not need a position on AGI to use these tools well, and a position is unlikely to help you. What helps is judging a system on the specific job you have for it: does it do this task, at what reliability, and what does a failure cost me? That question has an answer you can test this week, which no AGI timeline does.
Two habits carry over regardless of how the debate resolves. Check the output where being wrong matters, because a fluent answer and a correct one look identical and certainty in the wording tells you nothing about accuracy. And keep clear about what the thing in front of you is: software whose behavior comes from patterns in data, remarkable at that, and not currently anything else.
Common questions
Is ChatGPT AGI?
No, on any definition currently in use. Today's chat assistants are broad in subject matter and narrow in kind: they produce text, they cannot learn after release, and their reliability drops sharply on long tasks. Breadth of topic is easy to mistake for generality because language touches everything, but the underlying operation stays the same whatever you ask about.
How close are we to AGI?
Nobody knows, and anyone who sounds certain is usually answering a narrower question than the one you asked. The honest position is that progress on measurable tasks has been fast, several known gaps have not closed, and no one can say whether closing them needs more of the same or something new. Treat specific dates as arguments rather than forecasts.
What is the difference between AGI and superintelligence?
AGI is usually pitched at the level of a capable human across the board. Superintelligence means clearly beyond the best humans at essentially everything, including science and strategy. The two get discussed together because some people expect a short gap between them, on the argument that a system able to do research could improve itself. That expectation is contested and not evidence.
Will AGI take my job?
The question worth asking is narrower and more useful: which specific tasks in your job could a current system do at acceptable reliability, and which need judgment, accountability or physical presence. Task level change is already visible in drafting, summarizing and routine code. Whole occupation replacement runs into cost, regulation and liability, which move far more slowly than model capability.
Would we even know if AGI arrived?
Probably not at the time, and probably not all at once. With no agreed test, recognition would come from an accumulation of things working rather than a single announcement: systems finishing long projects unsupervised, picking up new tools without retraining, and being trusted with consequences. Expect the argument about whether it had happened to continue for years afterwards.