Symmetry at work

Every alias is a supercell

A cell that explains every line of a powder pattern is not a near miss and not a coincidence: its reciprocal grid contains the true one, which means its own cell is a superlattice of the true cell. So the ambiguity of indexing is the arithmetic of superlattices, and it can be counted — two cells with one unknown, sixteen with two, sixty-two with three, all of them accounting for the same twenty lines exactly.

Assumes Indexing a powder pattern, The figure of merit a supercell always beats and The sublattices that are the same shape.

Indexing a powder pattern solves the cubic case and says why it stops there: one unknown makes the arithmetic short enough to check completely, and it says in as many words that nothing there computes a lower-symmetry indexing.

The figure of merit a supercell always beats then establishes the reason a criterion is needed at all — a supercell’s grid contains the true cell’s, so the discrepancies are identical to every digit and no measure of agreement can separate them.

This essay is the arithmetic between those two. If an alias is always a supercell, then the aliases of a pattern are the superlattices of its cell, and how many there are is a question this collection has already learnt to answer.

Why an alias is a supercell and not a near miss

A powder pattern is a list of Q = 1/d². A candidate cell explains the list when every observed Q lies on the grid the cell generates.

Suppose it does. Then the true cell’s reciprocal lattice — which is where the observed Q values live — is contained in the candidate’s reciprocal lattice. Containment in reciprocal space reverses in direct space, so the candidate’s direct lattice is contained in the true one:

the candidate cell is a sublattice of the true cell, which is to say a supercell of it.

That is not an approximation and there is no tolerance in it. A cell that accounts for every line, exactly, is a supercell; a cell that is not a supercell fails on some line, and the two possibilities have nothing between them.

The alias accounts for every line and predicts more. The observed lines above, and below them the grid of a supercell that explains all of them. The full ticks are the observed lines, which the alias reproduces exactly; the faint ones are lines the alias predicts and nobody saw. That second set is the only thing that separates the two cells, and it is why an indexing criterion has to charge for unobserved lines rather than measure agreement.
Fig. 1 The observed lines above, and below them the grid of a supercell. The full ticks are the observed lines, which the alias reproduces exactly; the faint ticks are lines the alias predicts and nobody saw. Agreement is identical between the two cells and only the absences differ — which is the whole of why an indexing criterion must charge for unobserved lines.

Counting them

Once the aliases are superlattices they can be enumerated rather than searched for, and the enumeration is different for each crystal system because a supercell has to stay in the system.

For a cubic cell, the supercell must stay cubic. There is exactly one cubic sublattice at each of , 2m³ and 4m³, and only the first kind gives a primitive cubic cell, so the aliases are a, 2a, 3a, … — a single sparse series.

For a tetragonal cell, the square base may be thinned to any square sublattice and the axis to any multiple, and the two choices are independent. Which square sublattices exist is Fermat’s question: a square sublattice of a square lattice has index a sum of two squares. So the tetragonal aliases are a product — one factor from the base, one from the axis.

For an orthorhombic cell, each of the three axes is multiplied independently, and the aliases are a triple product.

Two, sixteen and sixty-two cells for the same twenty lines. How many cells below index twenty-four account for every line of a pattern their own cell generated. Each of them is checked against the list rather than assumed to fit, and every one accounts for all twenty to the last digit — so no measure of agreement can separate them, however carefully the lines are measured.
Fig. 2 The number of cells below index twenty-four that account for every line of a pattern their own cell generated. Each is checked against the list rather than assumed to fit, and every one accounts for all twenty lines to the last digit.

Two cells for one unknown, sixteen for two, sixty-two for three. That is the difficulty of indexing stated as a number, and it is worth being precise about what kind of difficulty it is. The search is not slow; the answer is not unique. No amount of care in the searching removes a cell that genuinely explains every line, and no improvement in the measurement removes it either, because the agreement is exact rather than close.

The true cell, and the fifteen that are just as good. The smallest aliases of a tetragonal cell with a = 4 and c = 6, each one checked against the whole line list. The first row is the truth; every other row thins the square base by an index that is a sum of two squares, multiplies the axis by a whole number, or both — and accounts for every observed line exactly, because its reciprocal grid contains the true one.
Fig. 3 The smallest aliases of a tetragonal cell with a = 4 and c = 6. The first row is the truth; every other thins the square base by an index that is a sum of two squares, multiplies the axis by a whole number, or both. All of them are checked against the whole line list.

How fast it gets worse

The count depends on how large a cell is allowed to be, and the shape of that dependence is the shape of the problem.

Each unknown multiplies the ambiguity. The number of cells that explain the pattern, against how large a cell is allowed to be. With one unknown the count is flat — the only aliases are whole multiples of the edge, and there are two below sixty-four. With three it climbs past ninety over the same range, because each axis contributes its own multiples and the aliases are a product.
Fig. 4 The alias count against the largest index allowed. With one unknown the curve is flat — two cells below sixty-four, and that is all there will ever be. With three it climbs past ninety over the same range, because each axis contributes its own multiples and the aliases multiply.

With one unknown the count barely moves: below index sixty-four there are two aliases, and allowing a bigger cell adds almost nothing, because a cubic supercell has index a perfect cube and the cubes are sparse. A cubic indexing is nearly unambiguous and it is a fact about the cubes.

With three unknowns the count is ninety-seven below the same bound. The indices available are all products i·j·k, and the number of ways to write an integer as an ordered triple grows quickly — so allowing a cell twice as large roughly doubles the candidates again.

That is the practical content of “low-symmetry indexing is hard”. It is not that a triclinic search has six parameters to sweep rather than one; it is that when the sweep finishes, the number of cells that fit perfectly is in the dozens, and choosing between them is a separate problem with a separate criterion.

The search finds Fermat by itself

The enumeration above assumed the answer: it built the superlattices and checked each one. That is the right order of work and it is also a way to miss a kind of alias nobody thought of, so the account needs a control.

The control is a blind grid over the parameters — a sweep that knows nothing about sublattices, tries a range of values for each unknown, and keeps whatever explains the lines. Run over a square base, it accepts a thinning by 1, 2, 4, 5, 8, 9, 10, 13, … and refuses 3, 6, 7, 11, 12, 14, 15, …

The search finds Fermat without being told about him. Which factors the blind grid will accept as a thinning of the square base. It accepts exactly the sums of two squares and refuses the rest — three, six, seven, eleven — and nothing in the search knows what a square sublattice is. The aliases of a powder pattern are the superlattices of its cell, so the arithmetic that decides which superlattices exist decides which aliases exist.
Fig. 5 Which factors the blind grid will accept as a thinning of the square base. It accepts exactly the sums of two squares and refuses the rest, and nothing in it knows what a square sublattice is — the arithmetic is a consequence of which cells explain the lines, not an assumption about which cells to try.

Those are the sums of two squares, and the blind grid produced them from nothing but the requirement that every observed line be accounted for. It also finds exactly the cells the sublattice enumeration predicts and no others, which is what makes the enumeration a description of the aliases rather than a list of the ones somebody looked for.

A number-theoretic fact turning up in a diffraction search is the sort of thing this collection is for. Fermat’s theorem about which primes are sums of two squares is from 1640 and is about integers; it decides, here, which wrong unit cells a crystallographer will find in the output of an indexing program.

The same difficulty, in a different ring

A tetragonal cell has two free parameters. So does a hexagonal one — a triangular net and a spacing — and running the identical enumeration on it says whether “two unknowns” is one situation or two.

It is one situation and two arithmetics. A triangular sublattice of a triangular lattice has index a Loeschian number a² + ab + b² rather than a sum of two squares, which is the same swap the plane’s similar sublattices make between the Gaussian and the Eisenstein integers. Both lists are infinite, both are about half the integers, and they barely overlap: 2, 5, 8, 10, 17, 18, 20 are available to a square base and not a triangular one; 3, 7, 12, 19, 21 the other way about.

Two unknowns twice, with two different arithmetics. A tetragonal cell and a hexagonal one both have two free parameters, and both are ambiguous — but by different rings. A square base is thinned by a sum of two squares and a triangular one by a Loeschian number, so the two lists of admissible indices barely overlap. The ambiguity belongs to the count of unknowns; which cells are the aliases belongs to the base.
Fig. 6 The admissible base indices for each of the two two-parameter systems, and how many cells each admits below index twenty-four. The difficulty is the same — two unknowns, an answer that is not unique — and the wrong answers are different, because which superlattices exist is decided by the base rather than by the count of parameters.

The consequence for a reader of an indexing list is small and worth having: a hexagonal candidate at a base index of three is an alias, and a tetragonal one at a base index of three does not exist. Which multiples to be suspicious of depends on the system, and the two lists are the two rings.

What a tolerance costs

Everything above is exact. A cell either carries every observed line or it does not, and the population of cells that do is the population of superlattices — a handful, arithmetically determined, and completely described.

A real pattern is measured, so a real search accepts a cell whose lines fall near the observed ones. That admits cells which are not superlattices at all, and the count of them is worth measuring rather than gesturing at.

A tolerance lets in hundreds that are not supercells. The same grid run at a range of tolerances. Below a thousandth the only cells that fit are the exact superlattices; a little above it the population is almost entirely cells with no arithmetic relation to the truth at all. The clean count of aliases is a statement about a noiseless pattern, and a real search is dominated by the other kind.
Fig. 7 The same grid at a range of tolerances, with the cells that fit split into those that are genuine superlattices and those that are not. Below a thousandth in Q the exact aliases are the whole population; a little above it they are one per cent of it.

The transition is sharp. At a tolerance of a ten-thousandth the search returns ten cells and every one is a superlattice. At three thousandths it returns a hundred and fifty-nine, of which a hundred and forty-nine have no arithmetic relation to the truth. At a hundredth it returns fifteen hundred and the ten exact ones are lost in them.

So the clean count of aliases is a floor rather than a description of the real problem. The superlattices are always there, they cannot be measured away, and they are what remains when the data are perfect. Everything else is a function of how well the lines were measured, and it dominates as soon as the measurement is anything short of exact.

That is a useful way to read an indexing program’s tolerance parameter. Tightening it removes the accidental candidates and cannot remove the arithmetic ones; loosening it buys robustness against a real pattern’s errors at the cost of a candidate list that is mostly noise. The two populations respond to it in completely different ways, and knowing which is which is the difference between tuning a search and hoping.

What separates the true cell

Nothing in the pattern’s positions does, and the previous rung explains why. What separates them is the lines an alias predicts and nobody observed.

That is visible in the first figure: the alias reproduces every observed line and adds a set of faint ones. A criterion that measures agreement scores the two cells identically; a criterion that divides by the number of possible lines charges the alias for its predictions, and de Wolff’s figure of merit is exactly that division.

So the three rungs fit together as one argument. Rung one says a cubic indexing is decidable because the arithmetic of one unknown is short. Rung three says agreement cannot be the criterion because a supercell’s agreement is identical. This rung says why both are true at once: the aliases are the supercells, their number is an arithmetic function of the symmetry, and it is small enough to check in the cubic case and not in any other.

How special the cubic case is

The three growth curves have three shapes and it is worth naming them, because the difference between “hard” and “easy” here is a difference in exponent rather than in constant.

Cubic. The aliases are the cells a·m, whose indices are the cubes . Below a bound V there are about V^{1/3} of them — below sixty-four there are four, below a thousand there are ten. The list grows so slowly that allowing an arbitrarily large cell barely enlarges it, which is why a cubic pattern is effectively indexed once the first candidate is found.

Tetragonal. The aliases are pairs (n, m) with n a sum of two squares and n·m ≤ V. About half the integers below a bound are sums of two squares in the loose sense that matters here, so the count is roughly proportional to V — sixteen below twenty-four, twenty-one below sixty-four. Doubling the allowed cell size roughly doubles the list.

Orthorhombic. The aliases are ordered triples with i·j·k ≤ V, which is the summatory divisor function twice over: about V log²V. Sixty-two below twenty-four and ninety-seven below sixty-four, and rising faster than either.

So the three cases are separated by a power of the bound rather than by an awkwardness of the search, and the cubic case is not merely the easiest — it is the only one whose ambiguity does not grow. That is the honest content of the first rung’s remark that one unknown is short enough to check completely: what is short is not the sweep, it is the answer.

What a single crystal buys

The contrast worth drawing is with a cell from a bag of spots, which indexes a single-crystal measurement rather than a powder.

A powder pattern gives the lengths of the reciprocal vectors and nothing else: every direction has been averaged away by the sample’s random orientations. So a candidate cell has only to reproduce a list of numbers, and any superlattice does.

A single crystal gives the vectors themselves. A supercell’s reciprocal lattice still contains the true one, so it still accounts for every observed spot — the aliases have not gone away — but the observed spots now have positions, and the extra points a supercell predicts sit at places a detector was looking at and found nothing. The absence is a measurement rather than an inference, and it is direct enough that the aliasing is usually settled by inspection.

That is the whole of the difference between the two problems, and it explains why a powder indexing needs a figure of merit and a single-crystal indexing mostly does not. The information the powder threw away is exactly the information that would have separated the candidates.

It also says what a powder measurement would need in order to be unambiguous, which is not more precision. It would need a reason to believe that a line predicted and not seen is genuinely absent rather than weak — and the intensities of a powder pattern depend on the structure, which is what the indexing was supposed to be the first step towards. The ambiguity is circular rather than technical, and the merit criterion is the standard way of cutting the circle rather than resolving it.

What this does not do

It does not index a real pattern. The lines here are synthesised from a known cell, exactly, with no peak overlap, no zero-point error, no impurity lines and no missing weak reflections. Every one of those makes the real problem harder in a way this computation does not model — in particular a tolerance means cells that are not supercells can also fit, which is a second and messier source of candidates on top of the exact ones counted here.

What the account has to refuse. The fourth and fifth rows are the pair that turns an enumeration into a result. A blind grid over the parameters — which knows nothing about sublattices — finds exactly the cells the sublattice enumeration predicts and rejects every base index that is not a sum of two squares. Without that control the enumeration would be a list of the aliases somebody thought to look for.
Fig. 8 The account run against what must fail it. The blind grid is the control that turns the enumeration into a result, and the last row is the check that a cell which is not a supercell is refused outright rather than scraping through on a tolerance.

And centring is a second kind of ambiguity, counted nowhere here. Every alias above is a supercell — a coarser lattice, whose grid contains the true one. A centred cell is the opposite: a finer lattice described in a larger conventional cell, whose extra points impose extinction rules that remove lines rather than adding them. A pattern indexed on a primitive cell can therefore also be indexed on a centred cell twice the size, with half its reflections declared systematically absent, and no line in the pattern objects. That ambiguity is about which reflections are allowed rather than about which are possible, so it is a different arithmetic — the absences a space group makes is where it belongs — and the count of superlattices here neither includes it nor is affected by it. A real candidate list carries both kinds at once.

And it does not cover the triclinic case. The three systems here have one, two and three free parameters with orthogonal axes; a monoclinic cell has four and a triclinic six, with angles among them, and the supercells then include shears as well as multiples. The count grows accordingly and is not computed here — what the three cases establish is the mechanism and its rate, not the number at the bottom of the symmetry ladder.

The shape of the whole difficulty

Putting the pieces in order gives a statement about powder indexing that is arithmetic all the way down.

A pattern is a list of numbers. A cell explains it exactly when the cell is a superlattice of the true one. The superlattices of a cell that stay in its crystal system are counted by the sublattice arithmetic of its base — the cubes for a cubic lattice, sums of two squares for a square base, Loeschian numbers for a triangular one, ordered triples for three independent axes. So the size of the candidate list is decided before any data are collected, by the symmetry alone.

Nothing in that chain is about the sample, the instrument or the search. A better diffractometer produces the same aliases; a cleverer algorithm finds them faster and does not remove them; a longer pattern with more lines does not help, because a superlattice explains every line however many there are. The only thing that shortens the list is a criterion that charges for absences, and the only thing that shortens it further is a measurement that recovers the directions the powder averaged away.

That is a satisfying place for a symmetry argument to end up. The question “how hard is it to index this pattern” sounds like a question about data quality, and it turns out to have an answer that depends on the crystal system and the size of cell one is willing to entertain, and on nothing else at all.

The one thing to carry

An indexing program’s output is a ranked list, and a reader of that list should know what the entries are. They are not near misses and they are not numerical noise: they are the superlattices of the answer, they fit exactly, and there are as many of them as the arithmetic of that crystal system allows.

Which means the useful question about a candidate is not “how well does it fit” — all of them fit perfectly — but “what is its index relative to the smallest cell that also fits”. A candidate at index one is the answer. A candidate at index four is the answer with a doubled base, and it will be there in every run, on every pattern, for reasons that have nothing to do with the sample.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

The objects this essay names

Each one links to every other essay that touches it.

Figure of meritIndexIndexingPowder patternQuadratic formReciprocal latticeSublatticeSupercell