Symmetry at work

Indexing a powder pattern

A powder pattern is a list of numbers and a cell is six. Getting the second from the first is the first step of every powder study and the one that fails — because the arithmetic has many answers, and choosing between them is a ranking rather than a deduction.

Assumes What a powder pattern loses and Systematic absences.

A powder pattern loses the orientation. Every crystallite is at a different angle, so the sharp spots of a single-crystal pattern smear into rings, and a reflection that was a point with three indices becomes a position on one axis with one number.

Indexing is that collapse, inverted. Given the list of spacings, find the cell. It is the first thing done with a powder pattern, everything downstream depends on it, and it is the step that fails — often enough that a substantial literature exists about programs that do it and about how to tell when they have not.

What a powder pattern loses. Every reflection of the structure, binned by spacing. Reflections whose reciprocal vectors have equal length arrive at the same place and add together, so the two-dimensional pattern collapses to one axis and the number under each peak is how many reflections it holds.
Fig. 1 The forward direction, from the essay that measures the loss: a two-dimensional pattern collapsed onto one axis, with reflections of equal spacing landing on the same line. Some coincidences are forced by symmetry and some are arithmetic accidents.

The forward problem, in the quantity that makes it easy

The useful quantity is not the spacing d but Q = 1/d², because for a cubic cell of edge a

Q(hkl) = (h² + k² + l²) / a²

so every observed value is an integer times a common unit. Indexing a cubic pattern is therefore the problem of finding one number — the unit — such that all the observed values are integer multiples of it, and of the right integers.

10 lines from an F cubic cell. The line list of a face-centred cubic cell with edge 5.64 Å, in the form indexing works with. Each line is one value of h² + k² + l², and only some values occur: the centring forbids the rest, which is the same extinction rule that decides a habit. Where two families of indices share a value they share a line, and no experiment at any resolution separates them — that is the powder's loss, and it is why a structure is refined against a profile rather than against integrated intensities.
Fig. 2 The line list of rock salt’s cell: face-centred cubic, edge 5.64 Å. Only some integers occur — 3, 4, 8, 11, 12, 16 — because the centring forbids the rest, and that pattern of gaps is what distinguishes a face-centred lattice from a primitive one.

Which integers occur is the systematic absences again. A primitive cubic lattice permits 1, 2, 3, 4, 5, 6, 8, …; a body-centred one only the even ones; a face-centred one 3, 4, 8, 11, 12, …. So the sequence of gaps is a signature of the centring, and reading it off is the classical way of indexing a cubic powder by hand.

And the coincidences are permanent. The line at 9 is {300} and {221} at once — two families of planes with the same spacing, related by no symmetry operation whatever. No resolution separates them, in any experiment, ever. That is the powder’s loss stated as an arithmetic fact rather than as a caveat.

12 lines from a P cubic cell. The line list of a face-centred cubic cell with edge 4.05 Å, in the form indexing works with. Each line is one value of h² + k² + l², and only some values occur: the centring forbids the rest, which is the same extinction rule that decides a habit. Where two families of indices share a value they share a line, and no experiment at any resolution separates them — that is the powder's loss, and it is why a structure is refined against a profile rather than against integrated intensities.
Fig. 3 A primitive cubic cell for comparison, at 4.05 Å. Every integer occurs, so the sequence has no gaps — and a pattern with gaps is a pattern from a centred lattice, which is the first thing a hand indexer reads off the list.

Reading the gaps is where a centring is found, and it is worth being precise about what that means. The gaps are not evidence about the point group and not evidence about the space group’s screws and glides; they are evidence about the lattice, and specifically about which centring translations it has. Reading a space group from its absences is the finer version of the same reading on a single-crystal pattern, where the zones can be examined one at a time; a powder gives only the union, which is why it identifies a centring and rarely a space group.

The inverse problem has many answers

7 cells explain the lines; one of them is right. A line list from a face-centred cubic cell of 5.64 Å, with a realistic error added, handed to a sweep over every cubic cell between 2 and 12 Å in all three centrings. 7 distinct cells explain every line within the tolerance, and each is a genuine solution rather than a numerical accident. The true cell comes top by de Wolff's figure of merit — the last Q over twice the mean discrepancy times the number of lines the candidate says should have been visible — which punishes a candidate for predicting lines nobody saw. That is the whole of what makes indexing decidable in practice: not the arithmetic, which has many answers, but a criterion for preferring one.
Fig. 4 The same ten lines, with a realistic measurement error added, handed to a sweep over every cubic cell between 2 and 12 Å in all three centrings. Seven distinct cells explain every line within the tolerance. Each is a genuine solution.

Seven cells, and the search was not told which was right. The true one comes top, but it is not alone, and the others are not numerical accidents: they are cells whose allowed integers happen to line up with the observed values to within the error. A primitive cell of the same edge explains the lines by assigning them different integers and predicting thirteen lines that were not seen; a body-centred cell of edge 7.976 Å does the same with different arithmetic.

The ranking is what makes indexing work, and the criterion is de Wolff’s figure of merit:

M = the last Q, divided by twice the mean discrepancy times the number of lines the candidate says should have been visible by then.

Two things are being weighed. The discrepancy punishes a cell that fits badly; the count of predicted lines punishes a cell that fits by being generous — a cell predicting a hundred possible lines will explain any twenty numbers, and the figure of merit says so.

What pg scatters. The diffraction pattern computed from the atom positions alone. Where a glide plane is present, alternate reflections along a row cancel exactly — and those missing spots are how the glide is identified in an experiment, since the glide itself is never seen.
Fig. 5 The same arithmetic in the plane, where it can be drawn: an operation with a half-translation extinguishes a row of reflections, and the pattern of what is left is the kind of signature a powder’s line list carries in one dimension.

The ambiguity that no measurement removes

A doubled cell explains everything and predicts too much. Doubling a cell's edge quarters the unit of Q, so every line that was an allowed integer is still one — the line at 3 becomes the line at 12 — and the doubled cell explains the data exactly whenever the true cell does. That is not a defect of the search; it is a property of the problem, and no improvement in the measurement touches it. What distinguishes the cells is how many lines each predicts that were not seen, and that is what a figure of merit weighs.
Fig. 6 Doubling the cell’s edge quarters the unit of Q, so every line that was an allowed integer is still one: the line at 3 becomes the line at 12. The doubled cell explains the data exactly whenever the true cell does, and the tripled cell after it.

A supercell always works, and no improvement in precision touches it. This is not a limitation of the search or of the instrument; it is a property of the problem. Any cell whose lattice is a sublattice of the true one reproduces every observed line, with extra lines predicted that were simply not observed — and “not observed” is exactly what a weak reflection looks like.

What removes the ambiguity is not evidence but convention: report the smallest cell that explains the data. That convention is doing real work and it is worth noticing that it is a convention, because it is the same move as choosing a reduced cell or a conventional setting. A crystallographic answer is very often a canonical representative rather than a unique one.

And a supercell is sometimes the truth. A superstructure — an ordering of atoms on a larger repeat — produces exactly this: weak extra lines on a doubled cell, easily missed, and the reflections a superlattice adds is the essay about them. So the smallest-cell convention is not merely a tie-breaker but an assumption that can be wrong, and the way it is wrong is that a structure gets solved on the subcell and looks nearly right.

More data does not give one answer

More lines do not give one cell; they give a wider margin. The same pattern truncated to three, four, six, eight and twelve lines, with the search run again on each. The number of cells that explain the data does not fall to one — a supercell always explains it, and so do several unrelated cells within the measurement's tolerance. What changes is the figure of merit's margin: with three lines the right cell is barely ahead of its rivals, and with twelve it is far ahead. Indexing is a ranking problem rather than a solving problem, which is why every indexing program reports a list with scores instead of an answer.
Fig. 7 The same pattern truncated to three, four, six, eight and twelve lines, with the search rerun on each. The number of cells explaining the data does not fall to one; what grows is the margin by which the right one leads.

The instinct is that more lines must eventually pin the cell down. They do not, and the reason is the supercell family: at any number of lines, the multiples explain everything. What twelve lines buy over three is that the true cell’s figure of merit pulls away from its rivals, so the ranking becomes decisive even though the list does not become short.

That is why every indexing program reports a list with scores rather than an answer, and why the practical rule of thumb — a figure of merit above about ten for twenty lines, and a clear gap to the next candidate — is a rule about margins. A pattern that indexes with three candidates within a few per cent of each other has not been indexed.

The coincidences, and why they are not the difficulty

The accidental overlaps deserve a note, because they are the powder’s most-quoted defect and they are not what makes indexing hard.

Two kinds of line overlap. A symmetry overlap is a set of reflections related by the point group — {100}, {010}, {001} in a cubic crystal — which have equal spacings for a reason and would be measured once even in a single-crystal experiment. An accidental overlap is two unrelated families landing on the same number, like {300} and {221}, and no experiment separates them.

Neither obstructs indexing. Indexing needs only the positions, and an overlap costs one line from the list rather than corrupting the ones that remain. What overlaps damage is the step after indexing: extracting an intensity for each reflection, which is what a structure determination needs, and which becomes an under-determined problem when several reflections share one measured intensity.

That is why the modern practice is the Rietveld method — refine a structural model against the whole profile rather than against extracted intensities — and why powder structure solution was rare before it. The line positions were always enough to index; the intensities were the difficulty, and they still are.

What breaks it in practice

The computation here is the clean case, and the failures worth knowing about are all in what makes a real pattern not the clean case.

An impurity phase. Two substances in the sample give two interleaved line lists, and any cell that explains the union of them is wrong. Indexing programs fail on this constantly, and the standard defence is to allow a few lines to go unexplained — which weakens the criterion exactly where it was doing the work.

A systematic error in the spacings. Sample displacement shifts every line in a way that is not a scale factor, so no cell fits well and the figure of merit collapses for every candidate. A pattern that indexes badly is often not a hard structure but a misaligned instrument, and the diagnosis is that every candidate is bad rather than that several are equally good.

A wrong first line. Indexing methods that work from the first few lines — and most of the classical ones do — are wrecked by an extra weak line at low angle from an impurity or from the sample holder, because it enters the arithmetic as a constraint that no true cell satisfies. The defence is to try each line as the first, which multiplies the work by the number of lines and is exactly what the automated programs do.

Low symmetry. The cubic case here has one unknown. A triclinic cell has six, the line list is the same length, and the search space is enormous — which is why powder indexing is routine for cubic and hard for triclinic, and why the successful programs use very different strategies for the two.

And the absences must be read correctly. Getting the centring wrong reassigns every integer, which changes the cell edge by a factor of √2 or 2. That is not a small error, and it is the reason a wrongly indexed pattern usually gives an edge related to the true one by a simple factor rather than a random number.

9 cells explain the lines; one of them is right. A line list from a face-centred cubic cell of 4.05 Å, with a realistic error added, handed to a sweep over every cubic cell between 2 and 12 Å in all three centrings. 9 distinct cells explain every line within the tolerance, and each is a genuine solution rather than a numerical accident. The true cell comes top by de Wolff's figure of merit — the last Q over twice the mean discrepancy times the number of lines the candidate says should have been visible — which punishes a candidate for predicting lines nobody saw. That is the whole of what makes indexing decidable in practice: not the arithmetic, which has many answers, but a criterion for preferring one.
Fig. 8 The same search on a primitive cell’s twelve lines. The winner is again the true cell, and the shape of the list is again a family of relatives rather than a set of near-misses — which is what an ambiguity of this kind looks like when it is drawn instead of described.

What the intensities do to the supercell

The supercell ambiguity is exact and no measurement of line positions removes it. The intensities are a different matter, and following what they say makes the ambiguity smaller than it first looks.

A doubled cell predicts every line the true cell predicts, and it also predicts lines between them. Those extra lines are observed to be absent — that is the whole reason the true cell was a candidate at all. So the doubled cell is a description in which a great many reflections happen to have zero intensity, and the question becomes why.

There are only two answers. Either the absences are systematic, in which case some space group on that larger cell forbids exactly that set of reflections and the description is legitimate; or they are not, in which case the doubled cell is asking for a coincidence among the structure factors of half its reflections, and no arrangement of atoms produces one.

That turns a convention into a test. Take the supercell, list the reflections it predicts and the pattern does not show, and ask whether that pattern of absences is the pattern any group produces. A doubling in one direction whose extra layers are all absent is a centring or a screw, and it is a real possibility to be kept. A doubling whose absent set matches nothing in the tables is refuted — not ranked lower, refuted — because it would require a structural coincidence of unbounded size.

What survives that test is exactly the interesting case, which is the superstructure. An ordered arrangement on a doubled cell produces the extra lines weakly rather than not at all, and the practical difficulty moves from ambiguity to sensitivity: whether a line at one part in a thousand of the strongest is present is a question about counting statistics and about the background, not about the arithmetic. So the honest position is that positions alone cannot choose the cell, intensities can refute most of the alternatives, and the ones they cannot refute are the ones worth looking at — which is a better outcome than the convention on its own, and it needs the group as well as the lattice.

Where the exactness stops

The line list is computed and then damaged on purpose. The spacings come from an exact cell, and a fixed pseudo-random relative error of three parts in ten thousand is added, which is a realistic laboratory number. Without it the figure of merit is meaningless — the discrepancies would sit at the last bit of a floating-point number and every candidate would score infinitely well — and with it the numbers behave the way real ones do.

The search is a sweep and finds everything. Stepping the cell edge over a range is the crudest possible method; it is used here precisely because it is exhaustive, so the ambiguities are visible rather than hidden behind a cleverer algorithm’s choice. A real program uses successive dichotomy or a Monte Carlo search over six parameters, and reports a ranked list for the same reason.

Only the cubic case is computed. The arithmetic of one unknown is short enough to check completely, which is the site’s usual reason for choosing a case. Nothing here computes a triclinic indexing, and no claim about the general problem is made beyond the qualitative one that the search space is larger.

And the figure of merit is a heuristic. M is a good criterion and it is not a proof of anything; a wrong cell can score well and a right cell can score badly on a poor pattern. It ranks candidates. Deciding is done afterwards, by whether the structure refines.

Who found it, and when

Hull and Debye and Scherrer invented the powder method independently in 1916 and 1917, and indexing was a hand computation from the start: Hull’s paper indexes iron’s pattern by trying integer ratios, which is exactly the calculation above.

The classical hand methods are graphical — Hull–Davey charts, Bunn charts, the Ito method — and each is a way of turning the search into a curve-matching exercise a person can do. They work for high symmetry and give out below orthorhombic.

P. M. de Wolff’s figure of merit dates from 1968, and it is the paper that made indexing a subject with a criterion rather than a subject with opinions. Its two ingredients are the ones above: how well the lines fit, and how many the cell would have allowed. A companion measure, Smith and Snyder’s F₃₀, weighs the same trade differently and is quoted alongside it.

The programs are named after their authors and their methods — ITO, TREOR, DICVOL, and later Monte Carlo searches — and the standard practice is to run several and see whether they agree. That is a sociological answer to a mathematical difficulty, and it is a reasonable one: distinct search strategies failing in the same way is much less likely than one strategy failing.

What the search must refuse

A search that always returns a cell would be worthless, and the check is the obvious one: hand it numbers that are not a powder pattern.

Seven invented spacings, chosen not to be an integer sequence, are indexed by nothing — no cubic cell in the range, in any of the three centrings, explains them within the tolerance. That is the result the whole essay rests on, because without it “seven cells explain the lines” would be a statement about the tolerance rather than about the data.

The two refusals bracket the method neatly. Real data admits several cells and the figure of merit ranks them; invented data admits none. The distance between those two outcomes is what an indexing program is measuring, and a program that returned an answer for the invented list would be one whose answers on real data meant nothing.

The supercell result is the third check and it runs the other way. There the machinery must accept — every multiple of the cell has to explain the lines, because it genuinely does, and a search that rejected them would be hiding a real ambiguity behind a convention. Two of the checks require acceptance and one requires refusal, which is the shape a measurement’s checks usually take once they are honest about what the method can and cannot decide.

Where this anchor starts

This is the first rung of a new ladder, and the ground it takes is the inverse direction. Everything else in this field on the site runs forwards: a group produces a pattern, a structure produces intensities, a lattice produces a line list. Indexing runs backwards from a measurement to a description, and it has the properties that direction usually has — many answers, a ranking rather than a proof, and a convention doing part of the work.

Two things it shares with the essays around it are worth stating. It is what an experiment cannot tell apart again, in a different place: there, two space groups with identical absences; here, a family of cells with identical lines. And the resolution is the same in both cases — the ambiguity is real, the arithmetic does not remove it, and what settles the answer is either more evidence of a different kind or an admitted convention.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Accidental overlapCentred latticeDecidabilityFigure of meritIndexingInterplanar spacingMeasurementPowder diffractionSystematic absenceUnit cell