Whether there is a centre is a statistic
Assumes The law that hides handedness and The phase problem.
Every other claim on this site is decidable. A pattern has a symmetry or it does not; the detector enumerates the candidate operations and keeps those that map the point set to itself; there is no tolerance to choose and no residual to interpret. This essay is the exception, and it earns a place because it is one.
The question is the one a crystallographer has to answer before choosing a space group: is the structure centrosymmetric? Diffraction cannot answer it directly. Friedel’s law makes every intensity distribution look centrosymmetric whether the structure is or not, and the systematic absences — which decide so much else — are silent on it, because a centre produces no absences.
What decides it is the shape of the distribution of the intensities, and the argument is A. J. C. Wilson’s, from 1949.
Two distributions, from one theorem
A structure factor is a sum of many terms with phases scattered around the circle. What happens to such a sum is the central limit theorem, and it happens differently depending on whether the terms come in pairs.
Without a centre, the real and imaginary parts of F are each sums of many independent contributions, so each is Gaussian and they are independent. The intensity |F|² is then the sum of two squared Gaussians — an exponential distribution, which has its mode at zero and a modest tail.
With a centre, the atoms come in pairs at ±r, every pair contributes 2f cos(2π h·r), and F is real. The intensity is the square of a single Gaussian: a χ² distribution with one degree of freedom, which piles up hard against zero and has a much heavier tail.
Those are two visibly different shapes for one measurable quantity, and the difference does not depend on the structure, only on whether the pairing is there. Normalising so that ⟨|E|²⟩ = 1 removes everything else — the number of atoms, their scattering power, the resolution — and leaves the two curves:
| acentric | centric | |
|---|---|---|
| ⟨|E|⟩ | 0.886 | 0.798 |
| ⟨|E|² − 1⟩ | 0.736 | 0.968 |
| ⟨|E|⁴⟩ | 2 | 3 |
The measurement
Two structures are built from a stated pseudo-random seed — the same atoms, with the centric one being those atoms plus their negatives, so the two differ in the property under test and in nothing else. Their intensities are computed by the same sum every diffraction figure on this site uses. The normalisation is a division by the mean, which is the only fitting anything here does.
The measured moments land where the theory says: 0.742 against 0.736 without a centre, 0.985 against 0.968 with one; ⟨|E|⟩ at 0.884 and 0.790 against 0.886 and 0.798. The agreement is at the level of the standard errors, and the standard errors are the point.
Reflections are counted in a half-plane, so that a Friedel pair contributes once. Counting both halves would halve the apparent variance for a bookkeeping reason and would make an acentric structure look centric — which is the kind of error that produces a wrong space group and no warning.
The number that should be quoted and usually is not
The test’s output is a decision, and a decision without a margin is not a measurement.
Two numbers are computed here for every structure. The margin is how much nearer the measurement is to one theory than to the other, in standard errors of its own mean. The misfit is how far it is from the theory it was assigned to.
For the honest pair, the margins are 10.5σ and 7.1σ and the misfits are 0.3σ and 0.5σ: the measurement is near one curve and far from the other, which is what a decisive test looks like. The separation between the two theories is about 34 standard errors at this sample size, which says the test could be decisive here even before it is run.
On a real data set neither number is that comfortable. A hundred reflections rather than twelve hundred multiplies the standard error by three and a half; absorption, scaling error and a partially-observed sphere all inflate it further; and a resolution range that includes the weak high-angle data is exactly where systematic error lives. The margin, not the verdict, is what says whether a test on real data has decided anything.
The case that fits neither curve
The classical failure of the test is a structure dominated by one heavy atom, and it is not an afterthought — it is the commonest structure a crystallographer meets, since a heavy atom is put in deliberately to solve the phase problem.
With one atom scattering ten times as strongly as the rest, the structure factor is no longer a sum of many comparable terms. The central limit theorem does not apply, the distribution is neither of the two, and the measured ⟨|E|² − 1⟩ comes out near 0.50 where the two curves say 0.74 and 0.97.
Read as a decision, that is “acentric” with a margin of twenty standard errors. Read properly, it is a misfit of twenty standard errors from the theory it was assigned to — a measurement that resembles the nearer of two curves while resembling neither. Both numbers are computed here for exactly that reason, and the second is the one that carries the information.
The refusal: one reflection decides nothing
Take any single reflection and ask which distribution it came from. Both distributions give every intensity positive probability, so the answer is never better than a guess dressed up.
Measured on a centrosymmetric structure, asking of each reflection which density is larger at this value gets the answer right for 43% of them — worse than a coin, because the exponential distribution is larger than the χ² over the middle of the range where most reflections sit.
That is the sense in which the property is not decidable in this site’s usual sense. It is not that the computation is hard; it is that no finite set of exact operations on the measured quantities returns it. The information is in the ensemble and is absent from every member of it.
What the test is used for
In practice this is one of three tools that between them fix a space group, and it is the one that answers what the other two cannot.
Systematic absences give the lattice centring and the screw and glide operations, by reading a space group from its absences. They leave the centre undecided, because a centre produces no absences at all.
The Laue class gives the point-group symmetry of the intensities, which is the true point group with a centre added — the eleven that contain inversion. It cannot say whether the added centre was already there.
The intensity statistics are the remaining half. The absences narrow the possibilities to a handful of groups; the statistics choose between the centrosymmetric ones and the rest. Pnma and Pna2₁ have the same absences and differ by a centre; the statistics are what separate them, and getting it wrong means refining a structure in a group that does not exist for it.
Reading the curve rather than the moments
The moments are a summary and the curve is the measurement, and where they disagree the curve is right.
The information is at small z. At z = 0.1 the acentric curve says 9.5% of reflections are weaker and the centric one says 25%; by z = 2 the two are within four points of each other. So the test is carried by the weak reflections — which are the ones measured worst, arrive with the largest relative errors, and are most often discarded as unobserved.
That is the practical difficulty in one sentence. The reflections that decide whether a structure has a centre are exactly the reflections an experiment measures least reliably, and a data set truncated at some signal-to-noise threshold has thrown away part of the evidence in a way that biases the answer towards acentric, since it removes weak reflections preferentially.
And it says why a centre makes weak reflections common. A real structure factor is a single number that must pass through zero to change sign, and it changes sign often; a complex one has to have both components vanish at once, which is a coincidence rather than a crossing. Half the reflections of a centrosymmetric structure have a negative sign, and the reflections near the crossings are the weak ones.
Where the exactness stops, precisely
The distributions are theorems and the measurements are samples. The two closed forms are exact consequences of the central limit theorem; every number measured against them has a standard error that falls as the square root of the number of reflections.
The normalisation is the one adjustable step. Dividing by the mean of |F|² is right when the mean is taken over a narrow enough resolution shell that the atomic scattering falloff is flat across it. On real data it is done shell by shell, and doing it globally biases the moments — an artefact that mimics a centre. Nothing here is measured on real data, so this essay’s numbers avoid the error rather than solving it.
A structure can be centrosymmetric and fail the test. A heavy atom does it, as above; so does a strong pseudo-translation, which correlates the reflections and breaks the independence the theorem needs. Both are common.
And the whole argument is about the structure, not about the crystal class. A centrosymmetric structure is one whose atom positions come in ± pairs about some origin. What the test reports is a property of that arrangement, and it happens to coincide with the presence of an inversion in the space group because the space group’s operations are what put the atoms in pairs.
The other thing that moves the same number
There is a second effect on the same statistic, it moves it in the opposite direction from a centre, and confusing the two is the commonest way this test is misread.
A merohedrally twinned crystal is one whose measurement is the sum of two orientations of the same structure, superposed reflection by reflection because the twin operation belongs to the lattice’s own symmetry. Each measured intensity is then a weighted sum of two intensities that are independent of one another, and summing two independent random quantities narrows the distribution of the result relative to its mean — the same central-limit effect that produced the two curves in the first place, applied once more.
So the moments move. Where an acentric structure gives ⟨|E|² − 1⟩ = 0.736 and a centric one gives 0.968, a perfectly twinned acentric structure gives about 0.5. That is not between the two curves; it is below both, and it is the reason the statistic is used as a twinning test at least as often as it is used as a centre test.
Read carelessly, though, it is a trap in both directions. A value of 0.75 might be an untwinned acentric structure, or a centric one that is partially twinned, and nothing in the number distinguishes them. A value near 0.5 is not evidence of an unusually acentric structure — there is no such thing — it is evidence that something has superposed two measurements, which is what a twin does.
The discipline that separates them is the same as everywhere else on this page: look at the curve rather than the moment. Twinning and centrosymmetry deform the cumulative distribution in different shapes, not merely to different summary values, and the difference is visible at small z where the information is. A single number computed from a distribution can be produced by many distributions, and reporting it without the curve behind it is reporting a projection and calling it the object.
It is also worth saying which order these run in. The scale and displacement factor come first, because the normalisation this whole test rests on needs them; the shape statistic comes second; and the twinning test is the same statistic read against a third hypothesis. Three questions, one distribution, and the arithmetic is shared — which is precisely why the answers get attributed to the wrong question.
Who found it, and when
A. J. C. Wilson published the distributions in Nature in 1949, in a short paper on the probability distribution of X-ray intensities, and followed it with the method for putting intensities on an absolute scale that also carries his name — the Wilson plot, which extracts the scale factor and the overall temperature factor from the same statistics.
Two things about the timing are worth noticing. The first is that this is a statistical argument arriving in a subject that had been resolutely deterministic; it belongs with the direct methods of Harker and Kasper (1948) and Sayre (1952), which are also statements about the distribution of structure factors rather than about any one of them. The second is that it made space-group determination routine at exactly the moment when structures started being solved in quantity.
The test has been refined since — Howells, Phillips and Rogers gave the cumulative N(z) form in 1950, which is the curve in the first figure and is more robust than the moments — and every modern structure-solution package runs some version of it automatically, reports it, and is ignored by users at their peril.
It is worth setting the two losses beside each other, because they are not the same kind of loss and the difference is what makes this essay’s subject a statistic rather than a measurement. The phases are gone outright: a detector records the square of a modulus, half the information in every reflection is discarded before anything is written down, and no amount of further measuring of the same kind returns it. Whether there is a centre is not gone at all. Every reflection carries a trace of it, the trace is dissolved across the whole data set rather than sitting in any part of it, and recovering it takes counting rather than an instrument. That is why the remedy for the first is to import information from somewhere else — a heavy atom, an anomalous scatterer, a positivity constraint — and the remedy for the second is simply to collect more reflections.
What a wrong answer costs
The test decides between a centrosymmetric group and a non-centrosymmetric one, and being wrong in each direction costs something different.
Choosing a centre that is not there forces every atom into a symmetric position it does not occupy. The refinement then has too few parameters, the residuals stay high, and the displacement parameters go strange — atoms acquire flattened or physically impossible ellipsoids as they try to represent two positions with one. It is uncomfortable and it is usually noticed.
Choosing no centre when there is one is the dangerous direction. The refinement then has roughly twice as many parameters as the data support, and the extra freedom is spent fitting noise. Everything looks better — the residual drops, the map is cleaner — and the coordinates acquire correlated errors that no statistic in the refinement reports. The classic symptom is a set of chemically implausible bond lengths that alternate long and short around a ring.
That asymmetry is why the intensity statistics are worth running before the refinement rather than after, and why a marginal result is a reason to try both groups rather than to pick the better-looking one.
The other statistic in the same measurement
The intensities carry a second piece of information that is worth naming here, because it is measured at the same time and is often confused with this one.
Wilson’s plot takes the mean intensity in each resolution shell and fits its logarithm against sin²θ/λ². The slope gives the overall temperature factor and the intercept gives the scale that puts the data on an absolute basis. That is a fit to the magnitude of the distribution; the centre test is a statement about its shape, and the two are independent.
They are usually run together and reported together, which is why they get conflated. The distinction matters in one specific way: the scaling has to be right before the shape means anything, since the normalisation divides by a mean that the scaling estimates. Getting the temperature factor wrong tilts the normalised intensities systematically with resolution, and the tilt shows up in the moments as a shift towards the centric values — a structure declared centrosymmetric because its data were scaled badly.
So the honest order is: scale, normalise shell by shell, then measure the shape. Doing the last step on globally normalised data is the commonest way to get a wrong answer from a correct test.
The shape of this rung is worth keeping. Everywhere else this collection decides a symmetry in integers; here the property is real, the structure either has the centre or it does not, and what an experiment records is a distribution rather than an answer.
Where the ladder goes next
This rung establishes that a symmetry can be a statistical property of a measurement rather than a decidable property of a structure. Three rungs sit above it.
Wilson’s other statistic, the plot that scales the data and extracts the temperature factor, which is the same distribution used for the absolute value of the mean rather than for its shape.
The intensity statistics of a pseudo-symmetric structure, where a near-symmetry that is not exact shows up as a distribution that is neither the centric nor the acentric one — the statistical face of near symmetry and the tolerance, and the case where a mistaken space group is most likely.
Direct methods, which take the statistical view seriously enough to phase a structure with it: the distribution of triple products of normalised structure factors is not uniform, and the departure from uniformity is enough to determine phases. That is the most consequential thing the statistical view of diffraction ever produced, and it starts here.
What this makes readable
Essays that name this one as a prerequisite.
What links here
The 8 essays that link to this one and share the most of its objects, of 16 that link here.
- A translation that is nearly there
- A map of the atoms that break the law
- How many reflections it takes to know there is a centre
- How much of it is the other hand
- One experiment gives the cosine, the other gives the sine
- The average that knows the atoms and not where they are
- The relation that can say no
- Three phases that do not move when the origin does
The objects this essay names
Each one links to every other essay that touches it.
Central limit theoremCentrosymmetricFriedels lawIntensity distributionNormalised structure factorSpace group determinationStandard errorWilson statistics