Expected goals has been in mainstream football coverage for long enough now that most people have a settled opinion about it, and the opinions tend to be quite strong in both directions. Either it's the most important analytical advance in the sport's history, or it's a number that tells you a shot that didn't go in should have.
Both camps are usually arguing about something other than what the metric actually does.
What it is, plainly
Take every shot ever recorded in a large database. For each one, note the things you know about it before it was struck: distance from goal, angle, whether it was a header or a foot, whether it came from a cross, whether it was a one-on-one, how many defenders were between the ball and the goal.
Now find all the historical shots that look like the one you're interested in, and check what proportion went in. If 12% of shots from that position and situation were scored, the shot has an xG of 0.12.
That's it. It's a historical conversion rate for shots of that type. It isn't a judgement about whether the shot should have gone in, and it isn't a prediction about that specific shot.
The most common misunderstanding
People treat a single shot's xG as a statement about that shot. It isn't. A 0.12 shot going in isn't a surprise or a fluke — twelve percent of them do go in, that's the whole meaning of the number.
The metric only becomes informative in aggregate. One shot tells you nothing. Ninety minutes of shots tells you a little. A season of shots tells you quite a lot, because at that volume the noise starts to cancel out and you can see whether a team is generating good chances or bad ones.
The single-match xG scoreline you see on television is, honestly, the weakest possible use of the metric, and it's the one most people encounter. That's largely why it has a bad reputation.
What it's genuinely good at
Three things, and they're valuable.
Predicting future results better than past results do. This is the strongest empirical case for xG and it's quite well established. A team that's been outshooting opponents in quality terms but losing games will usually start winning them. A team winning while being outchanced will usually stop. Over a half-season, xG difference predicts the next half-season better than actual goal difference does.
Separating chance creation from finishing. If a team creates 1.8 xG a match and scores 1.1 goals, that's a finishing problem, not a creation problem, and it calls for a different solution. Before xG this distinction was a matter of opinion.
Evaluating shot selection. A player taking twenty shots a game from thirty yards will have a lot of shots and very little xG. The metric makes visible something that was always true but hard to argue about.
Where it falls apart
Now the honest limitations, several of which are serious.
It doesn't know about the goalkeeper. Most public models don't include where the keeper was standing, which is obviously relevant. Some newer models incorporate post-shot information but that's a different metric measuring a different thing.
It doesn't know who took the shot. A model treats a shot from the penalty spot as identical whether it's taken by an elite finisher or a centre-back. Over a career, some players genuinely do outperform their xG consistently, which suggests finishing skill exists and the model can't see it.
It ignores everything that isn't a shot. A team that works the ball into the box repeatedly and can't get a shot away registers nothing. A move that ends with a defender making a last-ditch block never happened as far as xG is concerned. This is a genuine blind spot and it's why xG systematically underrates teams that face very deep defences.
Different providers give different numbers. Models vary in their inputs and their training data. The same match can have meaningfully different xG figures depending on who calculated it, which people forget when they quote a figure as though it were a measurement rather than an estimate.
Sample sizes are smaller than they feel. A team takes maybe 500 shots in a league season. That sounds like a lot until you're trying to distinguish a genuinely good finisher from a lucky one, at which point it's nowhere near enough. Most "player X overperforms his xG" claims are made on samples that can't support them.
How I'd actually use it
As one input among several, and mostly for questions about process rather than outcome.
Useful question: are we getting into good positions? xG per shot answers that well. Useful question: is our recent bad run about performance or luck? xG difference over ten matches gives you a decent steer.
Less useful question: was that a good shot to take? Sometimes a low-xG shot is correct because the alternative was worse. Less useful question: is this striker good? Far too many other things matter.
And I'd generally distrust any argument that rests entirely on xG without anybody having watched the football. The number is a summary of what happened, produced by a model with known blind spots, and it works best when it's checked against the thing it's summarising.
Which is a boring conclusion, but most useful conclusions about measurement are.