Every alias is a supercell
Assumes Indexing a powder pattern, The figure of merit a supercell always beats and The sublattices that are the same shape.
Indexing a powder pattern solves the cubic case and says why it stops there: one unknown makes the arithmetic short enough to check completely, and it says in as many words that nothing there computes a lower-symmetry indexing.
The figure of merit a supercell always beats then establishes the reason a criterion is needed at all — a supercell’s grid contains the true cell’s, so the discrepancies are identical to every digit and no measure of agreement can separate them.
This essay is the arithmetic between those two. If an alias is always a supercell, then the aliases of a pattern are the superlattices of its cell, and how many there are is a question this collection has already learnt to answer.
Why an alias is a supercell and not a near miss
A powder pattern is a list of Q = 1/d². A candidate cell explains the list when every observed Q lies on the grid the cell generates.
Suppose it does. Then the true cell’s reciprocal lattice — which is where the observed Q values live — is contained in the candidate’s reciprocal lattice. Containment in reciprocal space reverses in direct space, so the candidate’s direct lattice is contained in the true one:
the candidate cell is a sublattice of the true cell, which is to say a supercell of it.
That is not an approximation and there is no tolerance in it. A cell that accounts for every line, exactly, is a supercell; a cell that is not a supercell fails on some line, and the two possibilities have nothing between them.
Counting them
Once the aliases are superlattices they can be enumerated rather than searched for, and the enumeration is different for each crystal system because a supercell has to stay in the system.
For a cubic cell, the supercell must stay cubic. There is exactly one cubic sublattice at each of m³, 2m³ and 4m³, and only the first kind gives a primitive cubic cell, so the aliases are a, 2a, 3a, … — a single sparse series.
For a tetragonal cell, the square base may be thinned to any square sublattice and the axis to any multiple, and the two choices are independent. Which square sublattices exist is Fermat’s question: a square sublattice of a square lattice has index a sum of two squares. So the tetragonal aliases are a product — one factor from the base, one from the axis.
For an orthorhombic cell, each of the three axes is multiplied independently, and the aliases are a triple product.
Two cells for one unknown, sixteen for two, sixty-two for three. That is the difficulty of indexing stated as a number, and it is worth being precise about what kind of difficulty it is. The search is not slow; the answer is not unique. No amount of care in the searching removes a cell that genuinely explains every line, and no improvement in the measurement removes it either, because the agreement is exact rather than close.
How fast it gets worse
The count depends on how large a cell is allowed to be, and the shape of that dependence is the shape of the problem.
With one unknown the count barely moves: below index sixty-four there are two aliases, and allowing a bigger cell adds almost nothing, because a cubic supercell has index a perfect cube and the cubes are sparse. A cubic indexing is nearly unambiguous and it is a fact about the cubes.
With three unknowns the count is ninety-seven below the same bound. The indices available are all products i·j·k, and the number of ways to write an integer as an ordered triple grows quickly — so allowing a cell twice as large roughly doubles the candidates again.
That is the practical content of “low-symmetry indexing is hard”. It is not that a triclinic search has six parameters to sweep rather than one; it is that when the sweep finishes, the number of cells that fit perfectly is in the dozens, and choosing between them is a separate problem with a separate criterion.
The search finds Fermat by itself
The enumeration above assumed the answer: it built the superlattices and checked each one. That is the right order of work and it is also a way to miss a kind of alias nobody thought of, so the account needs a control.
The control is a blind grid over the parameters — a sweep that knows nothing about sublattices, tries a range of values for each unknown, and keeps whatever explains the lines. Run over a square base, it accepts a thinning by 1, 2, 4, 5, 8, 9, 10, 13, … and refuses 3, 6, 7, 11, 12, 14, 15, …
Those are the sums of two squares, and the blind grid produced them from nothing but the requirement that every observed line be accounted for. It also finds exactly the cells the sublattice enumeration predicts and no others, which is what makes the enumeration a description of the aliases rather than a list of the ones somebody looked for.
A number-theoretic fact turning up in a diffraction search is the sort of thing this collection is for. Fermat’s theorem about which primes are sums of two squares is from 1640 and is about integers; it decides, here, which wrong unit cells a crystallographer will find in the output of an indexing program.
The same difficulty, in a different ring
A tetragonal cell has two free parameters. So does a hexagonal one — a triangular net and a spacing — and running the identical enumeration on it says whether “two unknowns” is one situation or two.
It is one situation and two arithmetics. A triangular sublattice of a triangular lattice has index a Loeschian number a² + ab + b² rather than a sum of two squares, which is the same swap the plane’s similar sublattices make between the Gaussian and the Eisenstein integers. Both lists are infinite, both are about half the integers, and they barely overlap: 2, 5, 8, 10, 17, 18, 20 are available to a square base and not a triangular one; 3, 7, 12, 19, 21 the other way about.
The consequence for a reader of an indexing list is small and worth having: a hexagonal candidate at a base index of three is an alias, and a tetragonal one at a base index of three does not exist. Which multiples to be suspicious of depends on the system, and the two lists are the two rings.
What a tolerance costs
Everything above is exact. A cell either carries every observed line or it does not, and the population of cells that do is the population of superlattices — a handful, arithmetically determined, and completely described.
A real pattern is measured, so a real search accepts a cell whose lines fall near the observed ones. That admits cells which are not superlattices at all, and the count of them is worth measuring rather than gesturing at.
The transition is sharp. At a tolerance of a ten-thousandth the search returns ten cells and every one is a superlattice. At three thousandths it returns a hundred and fifty-nine, of which a hundred and forty-nine have no arithmetic relation to the truth. At a hundredth it returns fifteen hundred and the ten exact ones are lost in them.
So the clean count of aliases is a floor rather than a description of the real problem. The superlattices are always there, they cannot be measured away, and they are what remains when the data are perfect. Everything else is a function of how well the lines were measured, and it dominates as soon as the measurement is anything short of exact.
That is a useful way to read an indexing program’s tolerance parameter. Tightening it removes the accidental candidates and cannot remove the arithmetic ones; loosening it buys robustness against a real pattern’s errors at the cost of a candidate list that is mostly noise. The two populations respond to it in completely different ways, and knowing which is which is the difference between tuning a search and hoping.
What separates the true cell
Nothing in the pattern’s positions does, and the previous rung explains why. What separates them is the lines an alias predicts and nobody observed.
That is visible in the first figure: the alias reproduces every observed line and adds a set of faint ones. A criterion that measures agreement scores the two cells identically; a criterion that divides by the number of possible lines charges the alias for its predictions, and de Wolff’s figure of merit is exactly that division.
So the three rungs fit together as one argument. Rung one says a cubic indexing is decidable because the arithmetic of one unknown is short. Rung three says agreement cannot be the criterion because a supercell’s agreement is identical. This rung says why both are true at once: the aliases are the supercells, their number is an arithmetic function of the symmetry, and it is small enough to check in the cubic case and not in any other.
How special the cubic case is
The three growth curves have three shapes and it is worth naming them, because the difference between “hard” and “easy” here is a difference in exponent rather than in constant.
Cubic. The aliases are the cells a·m, whose indices are the cubes m³. Below a bound V there are about V^{1/3} of them — below sixty-four there are four, below a thousand there are ten. The list grows so slowly that allowing an arbitrarily large cell barely enlarges it, which is why a cubic pattern is effectively indexed once the first candidate is found.
Tetragonal. The aliases are pairs (n, m) with n a sum of two squares and n·m ≤ V. About half the integers below a bound are sums of two squares in the loose sense that matters here, so the count is roughly proportional to V — sixteen below twenty-four, twenty-one below sixty-four. Doubling the allowed cell size roughly doubles the list.
Orthorhombic. The aliases are ordered triples with i·j·k ≤ V, which is the summatory divisor function twice over: about V log²V. Sixty-two below twenty-four and ninety-seven below sixty-four, and rising faster than either.
So the three cases are separated by a power of the bound rather than by an awkwardness of the search, and the cubic case is not merely the easiest — it is the only one whose ambiguity does not grow. That is the honest content of the first rung’s remark that one unknown is short enough to check completely: what is short is not the sweep, it is the answer.
What a single crystal buys
The contrast worth drawing is with a cell from a bag of spots, which indexes a single-crystal measurement rather than a powder.
A powder pattern gives the lengths of the reciprocal vectors and nothing else: every direction has been averaged away by the sample’s random orientations. So a candidate cell has only to reproduce a list of numbers, and any superlattice does.
A single crystal gives the vectors themselves. A supercell’s reciprocal lattice still contains the true one, so it still accounts for every observed spot — the aliases have not gone away — but the observed spots now have positions, and the extra points a supercell predicts sit at places a detector was looking at and found nothing. The absence is a measurement rather than an inference, and it is direct enough that the aliasing is usually settled by inspection.
That is the whole of the difference between the two problems, and it explains why a powder indexing needs a figure of merit and a single-crystal indexing mostly does not. The information the powder threw away is exactly the information that would have separated the candidates.
It also says what a powder measurement would need in order to be unambiguous, which is not more precision. It would need a reason to believe that a line predicted and not seen is genuinely absent rather than weak — and the intensities of a powder pattern depend on the structure, which is what the indexing was supposed to be the first step towards. The ambiguity is circular rather than technical, and the merit criterion is the standard way of cutting the circle rather than resolving it.
What this does not do
It does not index a real pattern. The lines here are synthesised from a known cell, exactly, with no peak overlap, no zero-point error, no impurity lines and no missing weak reflections. Every one of those makes the real problem harder in a way this computation does not model — in particular a tolerance means cells that are not supercells can also fit, which is a second and messier source of candidates on top of the exact ones counted here.
And centring is a second kind of ambiguity, counted nowhere here. Every alias above is a supercell — a coarser lattice, whose grid contains the true one. A centred cell is the opposite: a finer lattice described in a larger conventional cell, whose extra points impose extinction rules that remove lines rather than adding them. A pattern indexed on a primitive cell can therefore also be indexed on a centred cell twice the size, with half its reflections declared systematically absent, and no line in the pattern objects. That ambiguity is about which reflections are allowed rather than about which are possible, so it is a different arithmetic — the absences a space group makes is where it belongs — and the count of superlattices here neither includes it nor is affected by it. A real candidate list carries both kinds at once.
And it does not cover the triclinic case. The three systems here have one, two and three free parameters with orthogonal axes; a monoclinic cell has four and a triclinic six, with angles among them, and the supercells then include shears as well as multiples. The count grows accordingly and is not computed here — what the three cases establish is the mechanism and its rate, not the number at the bottom of the symmetry ladder.
The shape of the whole difficulty
Putting the pieces in order gives a statement about powder indexing that is arithmetic all the way down.
A pattern is a list of numbers. A cell explains it exactly when the cell is a superlattice of the true one. The superlattices of a cell that stay in its crystal system are counted by the sublattice arithmetic of its base — the cubes for a cubic lattice, sums of two squares for a square base, Loeschian numbers for a triangular one, ordered triples for three independent axes. So the size of the candidate list is decided before any data are collected, by the symmetry alone.
Nothing in that chain is about the sample, the instrument or the search. A better diffractometer produces the same aliases; a cleverer algorithm finds them faster and does not remove them; a longer pattern with more lines does not help, because a superlattice explains every line however many there are. The only thing that shortens the list is a criterion that charges for absences, and the only thing that shortens it further is a measurement that recovers the directions the powder averaged away.
That is a satisfying place for a symmetry argument to end up. The question “how hard is it to index this pattern” sounds like a question about data quality, and it turns out to have an answer that depends on the crystal system and the size of cell one is willing to entertain, and on nothing else at all.
The one thing to carry
An indexing program’s output is a ranked list, and a reader of that list should know what the entries are. They are not near misses and they are not numerical noise: they are the superlattices of the answer, they fit exactly, and there are as many of them as the arithmetic of that crystal system allows.
Which means the useful question about a candidate is not “how well does it fit” — all of them fit perfectly — but “what is its index relative to the smallest cell that also fits”. A candidate at index one is the answer. A candidate at index four is the answer with a doubled base, and it will be there in every run, on every pattern, for reasons that have nothing to do with the sample.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- The shapes a lattice in space can thin to index · quadratic form · sublattice
- A bigger cell, and sometimes the mirror index · sublattice
- A lattice is not a subgroup index · sublattice
- A row written as a product index · sublattice
- A screw that contains its own mirror image index · sublattice
- An ideal across and a prime along index · sublattice
The objects this essay names
Each one links to every other essay that touches it.
Figure of meritIndexIndexingPowder patternQuadratic formReciprocal latticeSublatticeSupercell