The figure of merit a supercell always beats
Assumes Indexing a powder pattern and A cell from a bag of spots.
Indexing a powder pattern does not produce a cell. It produces a list of cells that account for the observed lines, ranked by a number, and the reason every indexing program in the world prints a table rather than an answer is that the arithmetic genuinely has many solutions. Some of them are unrelated cells that happen to fit within the measurement’s tolerance. One family of them is guaranteed: every multiple of the true cell explains every line the true cell explains, for ever, and no improvement in the experiment removes a single member of it.
The situation has an unusual shape for an inverse problem. It is not that the data are noisy and the answer is uncertain; it is that the data are consistent with infinitely many answers exactly, and the choice among them is made on a principle that is nowhere in the measurement. The principle is that the smallest cell is reported, and it is a good principle — a structure described in a cell four times too large is described with four times too many atoms, most of them related by translations the description does not admit to having. But it is a convention, and a convention has to be turned into a number before a program can apply it.
So the ranking is the whole of the method, and the number doing the ranking is where the content is. The obvious candidate — a measure of how closely the cell accounts for the data, which is what a refinement reports and what a scientist reaches for — turns out to be the one number that is precisely useless here, and the reason is worth the essay.
The grid inside the grid
The cubic case makes the argument visible in one line. Working in Q = 1/d², a cubic cell of edge a puts its allowed lines at
Q = n / a², with n = h² + k² + l²,
so the allowed Q values are a grid of spacing 1/a², thinned by whatever systematic absences the centring imposes and by the integers that are not sums of three squares. A cell of edge ma has grid spacing 1/(m²a²), which is m² times finer.
The point is not that the finer grid has more places to land. It is that the finer grid contains the coarser one exactly. The line that the true cell indexes as n is indexed by the supercell as m²n, at the identical Q, with the identical residual. Every observed line keeps the same neighbour it had; nothing moves.
The consequence is immediate and total. Any statistic computed from the residuals alone — their mean, their root mean square, a weighted sum, a likelihood — returns the same value for the true cell and for every multiple of it. Not a similar value. The same one. A criterion of that shape cannot rank the family, because the family is one point as far as it can see.
That is a stronger statement than the usual one, which is that a supercell “also fits”. It fits identically, and the identity is exact rather than a matter of tolerance.
It is worth being precise about what “no improvement in the experiment removes it” means, because it sounds like impatience and is a theorem. Improving the experiment means one of three things: measuring the existing lines more precisely, measuring more lines to higher angle, or measuring weaker lines. The first shrinks the discrepancies equally for every member of the family, since they are the same discrepancies. The second and third add lines, and every added line is at an allowed Q for the true cell and therefore at an allowed Q for every multiple. A supercell is never contradicted by data; it is only ever made less plausible, and plausibility is not something a residual measures.
What has to be paid for
If agreement cannot distinguish the cells then something outside agreement must, and there is only one thing available: what the cell predicted that was not there.
The true cell of a rock-salt pattern says nineteen lines should have been visible by the last one observed, and twenty were seen — nothing is missing. Its double says seventy-four should have been visible, and the fifty-four that were not are the price. Its sextuple says six hundred and sixty-one, of which six hundred and forty-one are absent without explanation.
De Wolff’s M₂₀ is exactly that division. The last observed Q, over twice the mean discrepancy, over the number of lines the candidate says should have appeared by then. The first two factors are the fit, which is constant down the family; the third is the count, which is not. The entire discriminating power of the criterion lives in the divisor, and the parts of it that look like statistics are along for the ride.
Smith and Snyder’s F_N is built differently — the discrepancy in 2θ rather than in Q, the number of observed lines rather than the last of them — and has the same divisor and therefore the same behaviour, which is the sort of agreement between differently-shaped statistics that is worth more than either alone.
There is a second reading of the divisor that makes it look less like a penalty and more like an accounting. A candidate cell is a model, and the number of lines it permits by the last observation is a count of the model’s opportunities to be wrong. A cell permitting nineteen lines where twenty were seen has almost no freedom: nearly every place it allows a line, a line arrived. A cell permitting six hundred and sixty-one where twenty were seen has enormous freedom, and its agreement with the twenty is worth correspondingly less. The merit is the fit divided by the freedom, which is the shape every criterion for choosing between models of different complexity eventually takes.
The penalty is a square, and the volume is a cube
Now a number that is easy to get wrong, and it is wrong in an interesting direction.
A cell m times as large has m³ times the volume, and the count of reciprocal lattice points below a fixed radius is proportional to the volume, so the obvious expectation is that the divisor grows as m³ and the merit falls as m⁻³. Fitting the measured counts on logarithmic axes gives an exponent of −1.98.
The missing factor of m is the powder’s own loss, arriving where nobody expects it. A powder line is not a reflection; it is every reflection sharing a spacing, collapsed into one. A figure of merit counts lines, because lines are what was observed. In the cubic system the lines are the distinct values of h² + k² + l², and there are far fewer of those than there are triples — the number of distinct values below N grows as N, not as N^{3/2}, while the number of triples grows as N^{3/2}. Since N scales as m², the line count scales as m² and the reflection count as m³.
And the coefficient in front of that count is a fact from number theory. Not every integer is a sum of three squares: 7 is not, nor 15, 23, 28, 31, 39. Legendre’s theorem says the exceptions are exactly the integers of the form 4ᵃ(8b + 7), and their density is one sixth. Counting the first four thousand integers here gives 3,335 representable, which is 83.38% against the theorem’s 83.33%, with every single omission of Legendre’s form.
So the strength of the criterion that decides whether a powder pattern has been indexed correctly is set, to within a fraction of a percent, by which integers are sums of three squares. That is not a metaphor; it is the divisor.
A large enough cell indexes noise
The other half of what a figure of merit is for only appears when the data have nothing behind them.
Twelve spacings were invented — increasing numbers with no cell, no lattice and no structure anywhere in their history — spread over the range a real pattern occupies, and offered to every cubic cell between three and a hundred and twenty ångströms.
A thousand three hundred and sixty-one cells account for all twelve within the tolerance a real measurement warrants. The smallest that does is 36.1 Å, which is an entirely ordinary size for a molecular crystal. And the best of them scores 13,987 on agreement alone, against 6,416 for the true cell of a genuine rock-salt pattern.
That is the sentence the whole essay exists to produce. On a criterion of fit, a cell with nothing behind it indexing numbers drawn from nothing beats the correct answer to a real experiment by a factor of two. Not marginally, not in an adversarial edge case — comfortably, on the first invented list tried.
The same cell scores 4.8 on de Wolff’s merit against the true cell’s 337.7, because it is being charged for the two thousand two hundred and fifty-six lines it says should have been seen.
How much data changes, and what it does not
The natural response to an ambiguity is more data, and it is worth measuring exactly what more data buys.
With three lines the search finds eight cells that account for them and the right one is not always on top. With twelve, the count is smaller and the right cell leads by a wide margin. What has improved is the margin, not the uniqueness: the supercell family is untouched, because it was never a shortage of data, and the unrelated cells that survive at three lines are eliminated because a fourth line lands where they said nothing would.
So the two ambiguities behave completely differently under more measurement, and it is useful to keep them apart. A coincidental cell is a statistical accident that more lines destroy. A supercell is a structural consequence of the arithmetic that no quantity of lines touches. The first is what a tolerance is fighting; the second is what a divisor is fighting; and a criterion that confuses them will be tuned in the wrong direction when it misbehaves.
The two populations, and where they sit
A criterion is used as a threshold. The folklore is that M₂₀ above ten means the indexing is right, and a threshold is only worth having if the two populations it separates are actually apart, so both were computed on the same footing.
Fifteen real patterns — five cell edges against three centrings — and five invented lists. On the merit the worst real pattern scores 453 and the best invented one 6.5, a factor of sixty-nine with a wide empty gap in the middle. On the fit the worst real pattern scores 5,443 and the best invented one 16,052, and the populations are not separated at all: they are in the wrong order.
One divisor moves a statistic from useless to decisive, and the divisor is not a refinement of the fit. It is a different quantity entirely — a count of predictions, not a measure of agreement — and the ranking works because it is there.
What the machinery had to refuse
Every claim above is stated in a form that could have come out the other way, and the checks are run rather than described.
Two of those are worth naming as the ones that would have been easy to leave out. The fourth demands that the machinery succeed at something undesirable: unless an invented list can be indexed by something, the entire criterion is guarding against a failure mode that does not occur, and a guard like that is decoration. The sixth demands that the simpler statistic fail, because a check that only confirms the elaborate statistic works has not shown that the elaboration was necessary.
Where the exactness stops
Computed here: the line list of a cubic cell with a realistic pseudo-random error; every cubic cell in a range that accounts for it; three figures of merit over each; the ladder of multiples up to six; the fitted exponent of the penalty; the density of sums of three squares to four thousand with every exception classified; and twenty scored patterns, fifteen real and five invented.
Cubic only, and the restriction is not cosmetic. In a triclinic cell there are six parameters instead of one, the search space is enormous, and the count of possible lines grows as the volume with no degeneracy to slow it — so the exponent found here is a cubic number and would be nearer three for a general cell. The argument is general: the fine grid still contains the coarse one in any system, so the fit is still constant along a supercell family. Only the size of the penalty changes.
The tolerance is a choice and everything depends on it. Whether a cell “explains” a line is decided by a threshold on the discrepancy, and the count of cells that pass, the smallest cell that indexes noise and the separation between populations all move when it moves. Choosing it too loosely makes everything index; too tightly and a real pattern with an uncorrected sample displacement fails. This is the same difficulty near-symmetry has, in the one place where the number cannot be avoided by working in integers.
With exact data, none of this exists. If the spacings carry no error the discrepancies are at the last bit of a double, the fit is infinite for the true cell and for every supercell, and every ambiguity looks equally perfect. A figure of merit is a statement about a measurement, and it is meaningless applied to arithmetic. The error used here is a fixed pseudo-random relative shift of three parts in ten thousand, so a figure drawn twice is the same figure, and it stands in for peak position, wavelength calibration and sample displacement without pretending to model any of them.
The invented lists are invented in one particular way, and a different way would give different numbers. They are increasing sequences with a fixed step distribution, which is what an arbitrary list of spacings looks like; a list generated from a non-cubic cell would be harder to index cubically and would flatter the criterion. So the thousand three hundred and sixty-one cells that accept twelve arbitrary numbers is a measurement of how permissive the cubic system is against unstructured input, and not an estimate of how often a real indexing goes wrong. What it establishes is the direction of the effect and its size, which is what the argument needs.
And a merit is not a probability. M₂₀ = 337 does not mean anything is 337 times more likely. There is no distribution behind it and no test being conducted; it is an ordering, and the threshold of ten is a convention that has survived because it works, in the sense that people who ignore it publish wrong cells. What is computed here is the ordering’s behaviour, not its calibration.
Who found it, and what they were afraid of
Pieter de Wolff proposed M₂₀ in 1968, in a two-page note, at a moment when automatic powder indexing programs had just become possible and were producing cells that were confidently wrong. His concern was precisely the one above: that a program searching a large space would find something, and that the something would agree beautifully with the data.
The construction he chose is telling. He could have made the criterion more statistical and did not; the divisor is a raw count of possible lines, with no weighting and no theory of errors, because the failure he was guarding against is not a statistical failure. It is a cell that is too big, and too big is a matter of arithmetic.
Smith and Snyder’s F_N followed in 1979 with the observed-line count in the numerator, and the two have coexisted since — not because either is better, but because a cell that scores well on both is harder to doubt than one that scores well on either. That is the same argument this collection makes about two computations sharing no intermediate, arrived at by practitioners who needed it rather than by anybody looking for it.
The persistence of the supercell ambiguity through all of it is the thing to carry away. Sixty years of indexing programs have not removed it and cannot: it is a property of the problem, the convention that the smallest cell is reported is a convention, and a figure of merit is what makes the convention operational.
Where the ladder goes next
Back, to the search this ranking sits on top of: indexing a powder pattern, where the sweep over candidate cells finds everything that works rather than choosing an answer.
Sideways, to the same ambiguity in a single-crystal experiment, where the data are richer and the sublattice problem survives anyway: a cell from a bag of spots, and every way down, and no way round, which counts the sublattices a lattice has and so counts the ambiguity exactly.
And into the number theory the divisor turned out to rest on: how many vectors of each length, where the same sums of squares are counted for their own sake, and the lengths do not name the lattice, which is the statement that even a complete list of them leaves something undecided.
What this makes readable
Essays that name this one as a prerequisite.
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
Figure of meritIndexingOverdeterminationPowder diffractionRankingSum of three squaresSupercellSystematic absenceTolerance