How it is known

Whether there is a centre is a statistic

Everything else on this site is decidable: a pattern has a symmetry or it does not, and the detector settles it in integers. Whether a structure has an inversion centre is not like that. No single reflection carries the answer — the distribution of all of them does.

Assumes The law that hides handedness and The phase problem.

Every other claim on this site is decidable. A pattern has a symmetry or it does not; the detector enumerates the candidate operations and keeps those that map the point set to itself; there is no tolerance to choose and no residual to interpret. This essay is the exception, and it earns a place because it is one.

The question is the one a crystallographer has to answer before choosing a space group: is the structure centrosymmetric? Diffraction cannot answer it directly. Friedel’s law makes every intensity distribution look centrosymmetric whether the structure is or not, and the systematic absences — which decide so much else — are silent on it, because a centre produces no absences.

What decides it is the shape of the distribution of the intensities, and the argument is A. J. C. Wilson’s, from 1949.

N(z): the fraction of reflections weaker than z. The cumulative distribution of normalised intensities, measured on two structures built from the same atoms — one with an inversion centre, one without — and drawn against the two closed forms, 1 − e^(−z) without a centre and erf(√(z/2)) with one. The curves are furthest apart at small z, which is the useful end: a centrosymmetric structure has far more nearly-absent reflections, because its structure factor is a single real number that can pass through zero rather than a complex one that rarely does.
Fig. 1 The fraction of reflections weaker than z, measured on two structures built from the same atoms — one with an inversion centre, one without — against the two closed forms. The curves are furthest apart at small z, where a centre piles up weak reflections.

Two distributions, from one theorem

A structure factor is a sum of many terms with phases scattered around the circle. What happens to such a sum is the central limit theorem, and it happens differently depending on whether the terms come in pairs.

Without a centre, the real and imaginary parts of F are each sums of many independent contributions, so each is Gaussian and they are independent. The intensity |F|² is then the sum of two squared Gaussians — an exponential distribution, which has its mode at zero and a modest tail.

With a centre, the atoms come in pairs at ±r, every pair contributes 2f cos(2π h·r), and F is real. The intensity is the square of a single Gaussian: a χ² distribution with one degree of freedom, which piles up hard against zero and has a much heavier tail.

Those are two visibly different shapes for one measurable quantity, and the difference does not depend on the structure, only on whether the pairing is there. Normalising so that ⟨|E|²⟩ = 1 removes everything else — the number of atoms, their scattering power, the resolution — and leaves the two curves:

acentric centric
⟨|E|⟩ 0.886 0.798
⟨|E|² − 1⟩ 0.736 0.968
⟨|E|⁴⟩ 2 3

The measurement

The moments, with their errors. The three moments of the normalised intensities for a structure with a centre and one without, each with the standard error of its own mean. The two theories differ by about 0.23 in ⟨|E|² − 1⟩ and the measurements have errors of a few thousandths, which is what makes the test decisive here. On a real data set with a hundred reflections and absorption errors it is a great deal less so, and the error bar is the honest part of the answer.
Fig. 2 The moments measured on both structures, each with the standard error of its own mean, against the theoretical values. The two theories differ by about 0.23 in ⟨|E|² − 1⟩ and the measurements have errors of a few thousandths.

Two structures are built from a stated pseudo-random seed — the same atoms, with the centric one being those atoms plus their negatives, so the two differ in the property under test and in nothing else. Their intensities are computed by the same sum every diffraction figure on this site uses. The normalisation is a division by the mean, which is the only fitting anything here does.

The measured moments land where the theory says: 0.742 against 0.736 without a centre, 0.985 against 0.968 with one; ⟨|E|⟩ at 0.884 and 0.790 against 0.886 and 0.798. The agreement is at the level of the standard errors, and the standard errors are the point.

Reflections are counted in a half-plane, so that a Friedel pair contributes once. Counting both halves would halve the apparent variance for a bookkeeping reason and would make an acentric structure look centric — which is the kind of error that produces a wrong space group and no warning.

The number that should be quoted and usually is not

The test’s output is a decision, and a decision without a margin is not a measurement.

Two numbers are computed here for every structure. The margin is how much nearer the measurement is to one theory than to the other, in standard errors of its own mean. The misfit is how far it is from the theory it was assigned to.

For the honest pair, the margins are 10.5σ and 7.1σ and the misfits are 0.3σ and 0.5σ: the measurement is near one curve and far from the other, which is what a decisive test looks like. The separation between the two theories is about 34 standard errors at this sample size, which says the test could be decisive here even before it is run.

On a real data set neither number is that comfortable. A hundred reflections rather than twelve hundred multiplies the standard error by three and a half; absorption, scaling error and a partially-observed sphere all inflate it further; and a resolution range that includes the weak high-angle data is exactly where systematic error lives. The margin, not the verdict, is what says whether a test on real data has decided anything.

The case that fits neither curve

The case the test cannot decide. The same measurement made on three structures, the third of which has one atom scattering ten times as strongly as the others. Its moments sit outside both theories rather than between them — the sum has stopped being a sum of many comparable terms, so neither distribution applies — and the verdict it gets is confident and meaningless. The number to read is the misfit, in standard errors, from the theory it was assigned to.
Fig. 3 The same measurement on three structures, the third of which has one atom scattering ten times as strongly as the others. Its moment sits outside both theories rather than between them, and the verdict it gets is confident and meaningless.

The classical failure of the test is a structure dominated by one heavy atom, and it is not an afterthought — it is the commonest structure a crystallographer meets, since a heavy atom is put in deliberately to solve the phase problem.

With one atom scattering ten times as strongly as the rest, the structure factor is no longer a sum of many comparable terms. The central limit theorem does not apply, the distribution is neither of the two, and the measured ⟨|E|² − 1⟩ comes out near 0.50 where the two curves say 0.74 and 0.97.

Read as a decision, that is “acentric” with a margin of twenty standard errors. Read properly, it is a misfit of twenty standard errors from the theory it was assigned to — a measurement that resembles the nearer of two curves while resembling neither. Both numbers are computed here for exactly that reason, and the second is the one that carries the information.

The refusal: one reflection decides nothing

Take any single reflection and ask which distribution it came from. Both distributions give every intensity positive probability, so the answer is never better than a guess dressed up.

Measured on a centrosymmetric structure, asking of each reflection which density is larger at this value gets the answer right for 43% of them — worse than a coin, because the exponential distribution is larger than the χ² over the middle of the range where most reflections sit.

One reflection decides nothing: 43 per cent, not fifty. The two probability densities a normalised intensity can be drawn from — e^(−z) without a centre and the χ² form with one — and the measurement that says a single reflection carries no verdict. Asking of each of the 1200 reflections of a centrosymmetric structure which density is larger at its own value gets the answer right for 519 of them, which is 43 per cent: worse than guessing. The reason is visible in the crossing at z = 0.19. Above it the acentric density is the larger one, and that is where most reflections of any structure sit, so a majority of a centric structure's own reflections individually resemble the wrong curve. The information is in the shape of the ensemble and is absent from every member of it — which is the sense in which this property, alone on this site, is not decidable by any finite set of exact operations on the measured quantities.
Fig. 4 The two densities, and the measurement that says a single reflection carries no verdict. The crossing is located rather than eyeballed: above it the acentric density is the larger one, and that is where most reflections of any structure sit — so a majority of a centric structure’s own reflections individually resemble the wrong curve. The count printed underneath is of that structure’s reflections, one at a time, and it is below a half.

That is the sense in which the property is not decidable in this site’s usual sense. It is not that the computation is hard; it is that no finite set of exact operations on the measured quantities returns it. The information is in the ensemble and is absent from every member of it.

Both patterns have a centre and only one structure does. The intensities of two structures built from the same atoms — one with an inversion centre and one without — drawn as discs whose size is the fourth root of the intensity. Both pictures are centrosymmetric, and exactly so: the intensity at a reflection and at its opposite agree to within a billionth of the strongest reflection in each, which is Friedel's law and is a consequence of the scattering factors being real rather than of anything about the structure. So the direct route is closed — no symmetry read off a diffraction pattern can say whether the arrangement has a centre. The two patterns are nevertheless different from one another, and what differs is the shape of the distribution of the intensities rather than any symmetry of their arrangement.
Fig. 5 Why the direct route is closed, drawn on the two structures this essay measures. Each disc is a reflection, its size the fourth root of the intensity. Both pictures are centrosymmetric, and exactly so — the intensity at a reflection and at its opposite agree to within a billionth of the strongest in each — because Friedel’s law follows from the scattering factors being real and not from anything about the arrangement. The two patterns are nevertheless different from one another, and what differs is the distribution of the intensities rather than any symmetry of their positions.

What the test is used for

In practice this is one of three tools that between them fix a space group, and it is the one that answers what the other two cannot.

Systematic absences give the lattice centring and the screw and glide operations, by reading a space group from its absences. They leave the centre undecided, because a centre produces no absences at all.

The Laue class gives the point-group symmetry of the intensities, which is the true point group with a centre added — the eleven that contain inversion. It cannot say whether the added centre was already there.

The intensity statistics are the remaining half. The absences narrow the possibilities to a handful of groups; the statistics choose between the centrosymmetric ones and the rest. Pnma and Pna2₁ have the same absences and differ by a centre; the statistics are what separate them, and getting it wrong means refining a structure in a group that does not exist for it.

The eleven Laue classes. Adjoining the inversion to each of the thirty-two crystal classes collapses them onto 11 groups. Friedel's law says a diffraction experiment sees the crystal and its inverse alike, so this — and not the crystal class — is what a diffraction pattern's symmetry reports. The highlighted symbol in each row is the class that is already its own Laue class, which is to say the centrosymmetric one.
Fig. 6 The eleven Laue classes: the thirty-two crystal classes as diffraction sees them, each with a centre added. Two classes differing only by inversion are one Laue class, which is the gap this statistic exists to fill.

Reading the curve rather than the moments

N(z): the fraction of reflections weaker than z. The cumulative distribution of normalised intensities, measured on two structures built from the same atoms — one with an inversion centre, one without — and drawn against the two closed forms, 1 − e^(−z) without a centre and erf(√(z/2)) with one. The curves are furthest apart at small z, which is the useful end: a centrosymmetric structure has far more nearly-absent reflections, because its structure factor is a single real number that can pass through zero rather than a complex one that rarely does.
Fig. 7 The same cumulative curves on a fifth of the data and fewer than half the atoms — about the sample a small-molecule data set gives after the weak high-angle reflections have been cut. The measured curves are visibly rougher and both still sit on their own theory; the separation is largest below z = 1, and the reason is worth reading off the picture, because a centrosymmetric structure has many more nearly-absent reflections than an acentric one.

The moments are a summary and the curve is the measurement, and where they disagree the curve is right.

The information is at small z. At z = 0.1 the acentric curve says 9.5% of reflections are weaker and the centric one says 25%; by z = 2 the two are within four points of each other. So the test is carried by the weak reflections — which are the ones measured worst, arrive with the largest relative errors, and are most often discarded as unobserved.

That is the practical difficulty in one sentence. The reflections that decide whether a structure has a centre are exactly the reflections an experiment measures least reliably, and a data set truncated at some signal-to-noise threshold has thrown away part of the evidence in a way that biases the answer towards acentric, since it removes weak reflections preferentially.

And it says why a centre makes weak reflections common. A real structure factor is a single number that must pass through zero to change sign, and it changes sign often; a complex one has to have both components vanish at once, which is a coincidence rather than a crossing. Half the reflections of a centrosymmetric structure have a negative sign, and the reflections near the crossings are the weak ones.

Where the exactness stops, precisely

The distributions are theorems and the measurements are samples. The two closed forms are exact consequences of the central limit theorem; every number measured against them has a standard error that falls as the square root of the number of reflections.

The normalisation is the one adjustable step. Dividing by the mean of |F|² is right when the mean is taken over a narrow enough resolution shell that the atomic scattering falloff is flat across it. On real data it is done shell by shell, and doing it globally biases the moments — an artefact that mimics a centre. Nothing here is measured on real data, so this essay’s numbers avoid the error rather than solving it.

A structure can be centrosymmetric and fail the test. A heavy atom does it, as above; so does a strong pseudo-translation, which correlates the reflections and breaks the independence the theorem needs. Both are common.

And the whole argument is about the structure, not about the crystal class. A centrosymmetric structure is one whose atom positions come in ± pairs about some origin. What the test reports is a property of that arrangement, and it happens to coincide with the presence of an inversion in the space group because the space group’s operations are what put the atoms in pairs.

The evidence is in the weak reflections, and they are the worst measured. The fraction of reflections weaker than each of six thresholds, for a structure with an inversion centre and one without, built from the same atoms. The gap between the two is largest at the bottom — below z = 0.1 the centric structure has 25 per cent of its reflections against 9 — and has almost closed by z = 2. So the test is carried entirely by the weak reflections, which are the ones an experiment measures worst, reports with the largest relative errors and most often discards as unobserved. A data set truncated at a signal-to-noise threshold has thrown away part of the evidence in a way that biases the answer towards acentric. And nothing here is an absence: the weakest reflection in either structure is small and positive, because a centre produces no extinction rule of any kind — which is the gap this statistic exists to fill.
Fig. 8 Weak, and not absent. The fraction of reflections below each of six thresholds, for the two structures. The gap is wide at the bottom and closed by z = 2, so the evidence is carried entirely by the weak reflections — the ones an experiment measures worst. And the weakest reflection in either structure is small and positive: a centre produces no extinction rule of any kind, which is precisely why a statistic is needed where an absence would have done.
A structure, and the vectors between its atoms. On the left, 4 atoms in a cell. On the right, every one of the 16 vectors between them, each drawn from a common origin: 13 distinct positions, with the 4-fold peak at the origin being each atom paired with itself. That right-hand picture is what a Patterson map shows, and it is the thing a diffraction experiment gives without phases. It has more peaks than the structure has atoms — n² against n — which is why interpreting one is hard, and why it is always symmetric about its centre.
Fig. 9 The map an experiment can always compute, and the one place a centre is unavoidable: the Patterson function is centrosymmetric for every structure, centred or not, which is the real-space form of Friedel’s law and the reason the question needs a statistic.

The other thing that moves the same number

There is a second effect on the same statistic, it moves it in the opposite direction from a centre, and confusing the two is the commonest way this test is misread.

A merohedrally twinned crystal is one whose measurement is the sum of two orientations of the same structure, superposed reflection by reflection because the twin operation belongs to the lattice’s own symmetry. Each measured intensity is then a weighted sum of two intensities that are independent of one another, and summing two independent random quantities narrows the distribution of the result relative to its mean — the same central-limit effect that produced the two curves in the first place, applied once more.

So the moments move. Where an acentric structure gives ⟨|E|² − 1⟩ = 0.736 and a centric one gives 0.968, a perfectly twinned acentric structure gives about 0.5. That is not between the two curves; it is below both, and it is the reason the statistic is used as a twinning test at least as often as it is used as a centre test.

Read carelessly, though, it is a trap in both directions. A value of 0.75 might be an untwinned acentric structure, or a centric one that is partially twinned, and nothing in the number distinguishes them. A value near 0.5 is not evidence of an unusually acentric structure — there is no such thing — it is evidence that something has superposed two measurements, which is what a twin does.

The discipline that separates them is the same as everywhere else on this page: look at the curve rather than the moment. Twinning and centrosymmetry deform the cumulative distribution in different shapes, not merely to different summary values, and the difference is visible at small z where the information is. A single number computed from a distribution can be produced by many distributions, and reporting it without the curve behind it is reporting a projection and calling it the object.

It is also worth saying which order these run in. The scale and displacement factor come first, because the normalisation this whole test rests on needs them; the shape statistic comes second; and the twinning test is the same statistic read against a third hypothesis. Three questions, one distribution, and the arithmetic is shared — which is precisely why the answers get attributed to the wrong question.

Who found it, and when

A. J. C. Wilson published the distributions in Nature in 1949, in a short paper on the probability distribution of X-ray intensities, and followed it with the method for putting intensities on an absolute scale that also carries his name — the Wilson plot, which extracts the scale factor and the overall temperature factor from the same statistics.

Two things about the timing are worth noticing. The first is that this is a statistical argument arriving in a subject that had been resolutely deterministic; it belongs with the direct methods of Harker and Kasper (1948) and Sayre (1952), which are also statements about the distribution of structure factors rather than about any one of them. The second is that it made space-group determination routine at exactly the moment when structures started being solved in quantity.

The test has been refined since — Howells, Phillips and Rogers gave the cumulative N(z) form in 1950, which is the curve in the first figure and is more robust than the moments — and every modern structure-solution package runs some version of it automatically, reports it, and is ignored by users at their peril.

It is worth setting the two losses beside each other, because they are not the same kind of loss and the difference is what makes this essay’s subject a statistic rather than a measurement. The phases are gone outright: a detector records the square of a modulus, half the information in every reflection is discarded before anything is written down, and no amount of further measuring of the same kind returns it. Whether there is a centre is not gone at all. Every reflection carries a trace of it, the trace is dissolved across the whole data set rather than sitting in any part of it, and recovering it takes counting rather than an instrument. That is why the remedy for the first is to import information from somewhere else — a heavy atom, an anomalous scatterer, a positivity constraint — and the remedy for the second is simply to collect more reflections.

What a wrong answer costs

The test decides between a centrosymmetric group and a non-centrosymmetric one, and being wrong in each direction costs something different.

Choosing a centre that is not there forces every atom into a symmetric position it does not occupy. The refinement then has too few parameters, the residuals stay high, and the displacement parameters go strange — atoms acquire flattened or physically impossible ellipsoids as they try to represent two positions with one. It is uncomfortable and it is usually noticed.

Choosing no centre when there is one is the dangerous direction. The refinement then has roughly twice as many parameters as the data support, and the extra freedom is spent fitting noise. Everything looks better — the residual drops, the map is cleaner — and the coordinates acquire correlated errors that no statistic in the refinement reports. The classic symptom is a set of chemically implausible bond lengths that alternate long and short around a ring.

That asymmetry is why the intensity statistics are worth running before the refinement rather than after, and why a marginal result is a reason to try both groups rather than to pick the better-looking one.

The other statistic in the same measurement

The intensities carry a second piece of information that is worth naming here, because it is measured at the same time and is often confused with this one.

Wilson’s plot takes the mean intensity in each resolution shell and fits its logarithm against sin²θ/λ². The slope gives the overall temperature factor and the intercept gives the scale that puts the data on an absolute basis. That is a fit to the magnitude of the distribution; the centre test is a statement about its shape, and the two are independent.

They are usually run together and reported together, which is why they get conflated. The distinction matters in one specific way: the scaling has to be right before the shape means anything, since the normalisation divides by a mean that the scaling estimates. Getting the temperature factor wrong tilts the normalised intensities systematically with resolution, and the tilt shows up in the moments as a shift towards the centric values — a structure declared centrosymmetric because its data were scaled badly.

So the honest order is: scale, normalise shell by shell, then measure the shape. Doing the last step on globally normalised data is the commonest way to get a wrong answer from a correct test.

The shape of this rung is worth keeping. Everywhere else this collection decides a symmetry in integers; here the property is real, the structure either has the centre or it does not, and what an experiment records is a distribution rather than an answer.

Where the ladder goes next

This rung establishes that a symmetry can be a statistical property of a measurement rather than a decidable property of a structure. Three rungs sit above it.

Wilson’s other statistic, the plot that scales the data and extracts the temperature factor, which is the same distribution used for the absolute value of the mean rather than for its shape.

The intensity statistics of a pseudo-symmetric structure, where a near-symmetry that is not exact shows up as a distribution that is neither the centric nor the acentric one — the statistical face of near symmetry and the tolerance, and the case where a mistaken space group is most likely.

Direct methods, which take the statistical view seriously enough to phase a structure with it: the distribution of triple products of normalised structure factors is not uniform, and the departure from uniformity is enough to determine phases. That is the most consequential thing the statistical view of diffraction ever produced, and it starts here.

What this makes readable

Essays that name this one as a prerequisite.

What links here

The 8 essays that link to this one and share the most of its objects, of 16 that link here.

The objects this essay names

Each one links to every other essay that touches it.

Central limit theoremCentrosymmetricFriedels lawIntensity distributionNormalised structure factorSpace group determinationStandard errorWilson statistics