How it is known

How many reflections it takes to know there is a centre

The test for a centre of symmetry compares one average of the intensities against two theoretical values a quarter apart. Whether that is a measurement depends on how many reflections went into the average, and the only honest way to find out is to run the test on structures whose answer is already known and count the mistakes.

Assumes Whether there is a centre is a statistic, The average that knows the atoms and not where they are and A twin hides in the statistics.

Whether there is a centre is a statistic settles that the presence of an inversion centre leaves a signature in the distribution of the intensities rather than in any one of them. Normalise the intensities so their mean is one, take ⟨|E|² − 1⟩, and compare it to the two values the distributions predict: 0.736 without a centre and 0.968 with one. On the structures measured there the test answers correctly by more than three standard errors, and the essay stops.

It stops one question early. Those two values are about a quarter apart. The scatter of the statistic falls as one over the square root of the number of reflections, so the test works when the separation beats the scatter and does not work when it does not — and nothing in the argument says which side of that line an actual experiment is on.

How many reflections is enough? It is a question about the measurement rather than about the crystal, and it has an answer only in the form a question about a measurement can have one: run the test many times on structures whose answer is already known, and count how often it is wrong.

How many reflections the centre test needs. The error rate of two tests for a centre of symmetry against the number of reflections used, measured on 60 centrosymmetric and 60 non-centrosymmetric structures at each point. The moment test — the one in every textbook, comparing ⟨|E|² − 1⟩ to its two theoretical values — reaches one error in twenty at 160 reflections and one in a hundred at 320. A likelihood ratio, which uses each reflection's own value instead of one average, reaches the same at 40 and 80. The gap is the price of summarising a distribution by its mean, and it is about a factor of four.
Fig. 1 The error rate against the number of reflections used, for two tests. Each point is sixty centrosymmetric structures and sixty without a centre, generated from stated seeds. The upper curve is the test every textbook gives; the lower one uses each reflection’s own value instead of one average of them.

Two errors, never one number

Before anything is counted the two mistakes have to be kept apart, because they are not the same mistake and crystallography does not treat them the same way.

A false acentric is a centrosymmetric structure reported as having no centre. A false centric is the opposite. Choosing the centrosymmetric space group for a structure that has no centre superimposes a real feature on its own mirror image and averages it away, which is a wrong answer that refines perfectly well; choosing the non-centrosymmetric one for a structure that does have a centre costs twice as many parameters and leaves a correlated mess, which announces itself. So the second error is recoverable and the first is not.

They are reported separately here throughout, and there is a blunt reason for insisting on it. A test that answers “acentric” whatever it is shown has an overall error rate of exactly one half and a perfect record on every acentric sample. Any summary that quotes one number for the performance of a test is a summary that cannot tell such a test from a real one.

The measured curves show the asymmetry plainly. At forty reflections the moment test misses a centre 22% of the time and invents one 13% of the time; the harmful error is the commoner one, throughout, at every count. That is not an artefact of the samples — it follows from the shapes of the two distributions, since the centric distribution has the longer tail and a small sample of it looks like the other more often than the other way round.

What the textbook test throws away

The moment test takes however many reflections were measured and reduces them to one number. That is its whole appeal and it is also its whole cost, and the cost can be measured rather than asserted.

The alternative uses every reflection’s own value. For a normalised intensity z the centric density is proportional to z^(−½) e^(−z/2) and the acentric one to e^(−z); multiply those over the reflections and compare, and by the Neyman–Pearson lemma no test does better. Both tests are run here on exactly the same samples, so the comparison is a count rather than a theorem.

The moment test reaches one error in twenty at 160 reflections. The likelihood ratio reaches it at 40. For one error in a hundred the numbers are 320 and 80. The textbook test needs about four times as much data to say the same thing with the same confidence.

Four is a modest factor and this is a modest complaint. The moment test is a line of arithmetic that can be done in the head from a table of intensities, it needs no assumption about the shape of the distribution beyond its two theoretical means, and at the resolution a real experiment reaches — thousands of reflections, not hundreds — both tests are far past the point where either is in doubt. The factor of four matters exactly where the data is thin: a small crystal, a rapid screen, a shell of high-angle reflections examined on its own.

The likelihood ratio has one practical detail that has to be right or it silently stops being a test. The product of a few hundred densities underflows to zero, and a comparison of two zeros answers whichever way the floating point happens to break. It is summed as logarithms here for that reason, which is a statement about arithmetic rather than about crystallography and belongs in the same category as everything else this collection insists on computing rather than quoting.

Which reflections, not just how many

A detail of the sampling turned out to matter more than expected, and it is worth stating because the wrong version is the obvious one.

To ask what a test does with n reflections, take n reflections. The obvious way to take them is the first n the loop produces — and the loop produces them in order of index, so the first n are the innermost shell. A low-resolution shell is not a random sample of reflections. Its statistics are its own: few reflections, each dominated by the gross shape of the electron density rather than by the atoms, and a moment that differs measurably from the whole pattern’s.

The measurement of that: the innermost 84 reflections of one structure give ⟨|E|² − 1⟩ = 0.752 where a spread sample of the same size gives 0.888. The first number is near the acentric theory and the second is between the two, and they are the same structure. Everything above therefore thins the reflections by a stated shuffle rather than by truncation, and the shuffle is part of the reported method rather than an implementation detail.

This is the same hazard Wilson statistics exists to handle from the other side. There the point is that ⟨|F|²⟩ falls with angle and the normalisation has to be shell by shell; here the point is that a shell is a biased sample of the distribution as well as of the magnitudes.

What a power curve must not claim. Four negative tests. A test that always answers the same way must not look powerful; a single reflection must not look like evidence; the innermost reflections must not be used as a sample of reflections, because a low-resolution shell has statistics of its own; and a structure the test cannot decide must not come back decided.
Fig. 2 The negative tests. Each is a way this measurement could be wrong while every curve it produced still looked reasonable: a constant answer reported as a result, one reflection treated as evidence, the innermost reflections used as a sample of reflections, and a structure the test cannot decide coming back decided.

The structure where more data makes it worse

Everything so far is about noise, and noise is the easy half. A statistic with scatter around the right value gets better with more of it, and the curves above are what that looks like.

The interesting failure is the other kind, and this collection has met it before: the heavy atom. Wilson’s derivation needs the structure factor to be a sum of many comparable terms, so that the central limit theorem applies and the two densities are the ones written above. One atom with ten times the scattering power of the rest breaks that. The sum is then one large term plus a fluctuation, its distribution is neither of the two curves, and the moment lands somewhere that has nothing to do with whether there is a centre.

Take a structure that genuinely has a centre and give it one heavy atom.

Where more data makes it worse. Two error curves for the same test. On structures of comparable atoms the rate falls to zero, which is what a working test does. On a centrosymmetric structure with one atom 10 times heavier than the rest it climbs, from about two-thirds at ten reflections to one at twelve hundred, because the sum for a structure factor has stopped being a sum of many comparable terms and the statistic settles on the wrong side of the dividing line. More reflections do not fix that; they make the wrong answer certain. Nothing inside the test distinguishes this from the other curve.
Fig. 3 Two error curves for the same test. On structures of comparable atoms the rate falls to zero — a working test, getting better with data. On a centrosymmetric structure with one atom ten times heavier than the rest it climbs, from about two-thirds at ten reflections to one at twelve hundred.

The error rate goes to one. Not to a plateau, and not to chance: to certainty about the wrong answer. ⟨|E|² − 1⟩ settles at 0.810 for this structure, which is between the two theoretical values and nearer the acentric one, so the test calls it acentric — and more reflections do not move the statistic towards either curve, they only pin it more tightly to 0.810. Every additional reflection makes the wrong verdict more confident.

That is the difference between a test that is noisy and a test that is inapplicable, and the sharp form of it is this: nothing inside the test distinguishes the two curves. A crystallographer with the heavy-atom structure and twelve hundred reflections sees a moment far from both theories and a verdict with an enormous margin. The margin is real and it is a margin on the wrong quantity.

The existing essay reports the misfit alongside the verdict for exactly this reason — how far the measurement is from the theory it was assigned to — and the power curve is what says why that second number is not optional. A margin measures how firmly the test has chosen between two hypotheses. A misfit measures whether either hypothesis was ever the right one.

What a single reflection says, and what a hundred do

It is worth looking at the two distributions themselves once more, because the power curve is a statement about them and nothing else.

N(z): the fraction of reflections weaker than z. The cumulative distribution of normalised intensities, measured on two structures built from the same atoms — one with an inversion centre, one without — and drawn against the two closed forms, 1 − e^(−z) without a centre and erf(√(z/2)) with one. The curves are furthest apart at small z, which is the useful end: a centrosymmetric structure has far more nearly-absent reflections, because its structure factor is a single real number that can pass through zero rather than a complex one that rarely does.
Fig. 4 The two distributions the whole test rests on: the acentric one, which is an exponential, and the centric one, which has a pile-up at small intensities and a longer tail. Both give every value positive probability, which is precisely why one reflection decides nothing and why the question is how many.

The overlap of those two curves is the whole difficulty. Take any single reflection and ask which distribution it came from: both assign its value a positive density, so the best possible answer is a guess weighted by the ratio of the two densities — and that guess is wrong about half the time. The measured curve says so directly: at one reflection the moment test is wrong 50% of the time, which is a coin.

What changes with a hundred reflections is not that any one of them became informative. It is that the sample acquires a shape, and the shape is what the two hypotheses disagree about. This is the point at which a distribution stops being a summary of a population and becomes a measurement of it, and the number of reflections at which that happens is the number this essay is about.

The moments, with their errors. The three moments of the normalised intensities for a structure with a centre and one without, each with the standard error of its own mean. The two theories differ by about 0.23 in ⟨|E|² − 1⟩ and the measurements have errors of a few thousandths, which is what makes the test decisive here. On a real data set with a hundred reflections and absorption errors it is a great deal less so, and the error bar is the honest part of the answer.
Fig. 5 The three moments the test is usually stated in, for both hypotheses and for a measured structure. Each of them separates the two distributions and each has its own scatter, so each has its own power curve; the one drawn above is ⟨|E|² − 1⟩ because it is the one in general use, and the argument would run the same way for any of them.

Why the numbers are what they are

The moment test’s error rate should fall as the separation of the two theories divided by the standard error of the statistic, and that ratio grows as the square root of the number of reflections. So the count needed for a given error rate should be proportional to the inverse square of the separation — and it is, well enough to be worth checking rather than assumed.

The separation is 0.968 − 0.736 = 0.232. The standard deviation of |E|² − 1 is about 0.9 for the acentric distribution, so the standard error at n reflections is about 0.9/√n, and the two theories are one standard error apart at n ≈ 15. A test that fails half the time when its two hypotheses are one error apart and rarely when they are three apart puts the useful region around n = 130, which is where the measured curve crosses one in twenty.

The agreement is worth having because it says the measurement is measuring what it claims. It also says what the answer depends on: not on the crystal, not on the space group, and not on the resolution, but on one ratio — a fixed separation against a scatter that shrinks like √n. Every property of the curve above follows from that, including the fact that the last few points are all exactly zero rather than small: sixty samples cannot resolve an error rate below about one in sixty, and the curve stops being informative before the test does.

Where the exactness stops

Computed here: for each of eight reflection counts, sixty structures with a centre and sixty without, generated from stated seeds; every reflection in a window of thirty, normalised to unit mean and thinned by a stated shuffle; two verdicts per sample; and the error rates counted separately for the two kinds of mistake.

Sixty samples per point is the resolution of the whole measurement. An error rate of zero in the table means “below about one in sixty”, not zero, and the counts reported as thresholds are the first grid point at which the rate falls below the level rather than the exact crossing. Both are stated as such above rather than quoted as though they were sharp.

These structures are not crystals. Atoms are placed at random in a square cell with equal scattering power, there is no thermal motion, no absorption, no measurement error and no systematic absence. Every one of those makes a real experiment’s statistics worse, so the counts here are a floor: a real determination needs at least this many reflections and probably several times more.

The case the test cannot decide. The same measurement made on three structures, the third of which has one atom scattering ten times as strongly as the others. Its moments sit outside both theories rather than between them — the sum has stopped being a sum of many comparable terms, so neither distribution applies — and the verdict it gets is confident and meaningless. The number to read is the misfit, in standard errors, from the theory it was assigned to.
Fig. 6 The heavy-atom structure examined directly rather than through its error rate: its intensity distribution against the two theories. It resembles neither, which is the honest statement, and a test that must choose between two curves will choose one anyway.

And the test decides a distribution, not a space group. A centre of symmetry is one bit of the answer. Which of the eleven Laue classes a pattern shows is a separate question with its own evidence, and systematic absences are what narrow the Laue class to a space group. The statistic here contributes exactly one bit and it is worth being precise about which one.

Who worked it out

Wilson set out the two distributions in 1949, in the paper that also gave crystallography the plot that carries his name. The tests built on them — the N(z) cumulative curve, the moment ⟨|E|² − 1⟩, and the several other moments that separate the same two hypotheses — were standard within a decade and are in every structure-determination package.

The question of how much data they need was answered informally and by experience: enough reflections, in practice a few hundred, and everyone knew that heavy atoms spoil it. The formal statement is the Neyman–Pearson lemma of 1933, which predates the crystallography and says exactly which test is best; the measurement above is that theorem meeting this particular pair of distributions and reporting the size of the gap.

The heavy-atom caveat is older than either. It is in Wilson’s own papers as a warning about the applicability of the central limit theorem, and its modern form is a distinction between a structure whose statistics are hypercentric — several heavy atoms in special positions — and one that is merely dominated by one. Both break the test in the same direction and only one of them is usually noticed.

Where the ladder goes next

Back, to the test itself and the two distributions it compares: whether there is a centre is a statistic, where the moment and the cumulative curve are derived and the heavy-atom failure first appears.

Sideways, to the other things the same intensities are asked. A twin hides in the statistics recovers a twin fraction from the second moment, and a translation that is nearly there splits the reflections into populations that make the centre test answer yes about a structure that has none — which is the same failure as the heavy atom, from a different cause, and would show the same climbing curve.

And to what the normalisation itself rests on: the average that knows the atoms and not where they are, whose shell-by-shell fit is what makes ⟨|E|²⟩ equal to one in the first place.

The test is a prior, not a verdict

Everything measured here is a property of a statistic computed from intensities alone, and no crystallographer decides a space group that way. Saying what the test is for keeps the error rates in proportion.

The statistic uses no model. It reads the distribution of the normalised intensities and nothing else — not the chemistry, not the positions, not whether a sensible structure exists in either group. That is its strength, since it is available before any structure is solved, and it is why its error rates are what they are.

A refinement uses everything. Solve the structure in the centrosymmetric group and in its non-centrosymmetric subgroup, refine both, and compare: the residuals, the displacement parameters, the geometry. A structure refined in too low a symmetry has correlated parameters, unreasonable ellipsoids and bond lengths that scatter without cause, and one refined in too high a symmetry has a disordered atom where the extra operation forced two positions together.

So the statistic decides where to start, not what to publish. An error rate of one in twenty at a few hundred reflections is a good enough guide for choosing which model to try first, and a poor basis for a claim about the crystal — which is why the second model is tried at all.

And the automated checks run in the other direction. Software given a solved non-centrosymmetric structure searches for symmetry the model has and the space group does not, by looking for operations that map the atoms onto themselves within a tolerance. That test uses the positions rather than the intensities, and it catches the case this essay’s statistic gets wrong most often — a structure whose heavy atoms sit centrosymmetrically while its light atoms do not.

One data set, several tests

There is a statistical hazard in the way these tests are used in practice, and it is worth naming because nothing in a single power curve reveals it.

The same intensities are tested repeatedly. Whether there is a centre, whether the crystal is twinned, whether there is a pseudo-translation, whether the Laue class is the one assumed — each is a test on one set of numbers, and each is run as a matter of routine.

Error rates do not stay put under repetition. A test with a one-in-twenty false-positive rate, run on five independent questions, flags something spurious about a quarter of the time. The tests here are not independent, which changes the arithmetic and does not remove the effect: a pseudo-translation biases the intensity distribution in the direction the centre test reads as a centre, and twinning biases it the other way.

Which is the practical reason the tests are reported together rather than in sequence. A structure flagged by one test alone is a structure with a suspicious statistic; a structure flagged by several in a consistent direction has something structural going on — and telling those two apart is what a report of the individual statistics allows and a single verdict does not.