What a lattice forbids

Better counting, and the wrong model

A smooth error across reciprocal space, as large as the random one, hardly changes which disorder model the data choose: it enters both models' residuals and cancels from their difference. What it changes is where better data lead. With random error alone, counting harder decides more pairs correctly. With the smooth error left in, counting harder stops helping and starts deciding pairs for the model that did not make the data — at the 0.1% level, and more of them the better the counting.

Assumes What refinement takes back from a wrong model, The molecule size that hides a disorder and Every reflection, several times over.

What refinement takes back from a wrong model refined each of two disorder models against data made from the other. It found that the likelihood ratio, the difference of their residual sums in units of the error variance, still chose the right model for fifteen of eighteen pairs at sixty-four atoms, with a random error of three per cent. The difference between the two refined residuals was three tenths of a per cent, and summed over four hundred and sixty reflections it was decisive. That essay then drew a conclusion from a number it had not measured. Real data carry systematic errors of a few per cent, and “a three per cent systematic error is exactly the size of difference the likelihood ratio relies on”. So, it said, the molecule size at which the data stop choosing is set by the systematic error and not by the counting.

This essay measures it. The random error is replaced, or joined, by a smooth error of exactly the same size: a scale that varies slowly across reciprocal space, as an uncorrected absorption does. Both models are refined again, and the decisions are counted again.

The prediction was half wrong, and the half that was right has a different shape than expected. A smooth error the size of the random one hardly changes the decisions. At sixty-four atoms it leaves fourteen of the eighteen pairs decided for the right model, against fifteen, because it enters both models’ residuals and cancels almost entirely from their difference. What it changes is where better data lead. With random error alone, measuring more carefully decides more pairs correctly. With the smooth error left in, measuring more carefully stops helping at fourteen or fifteen pairs. Beyond that point it starts deciding pairs for the model that did not make the data, and the better the counting, the more of them. The smooth error does not blur the verdict. It gives it a sign, and better counting turns that sign into a confident wrong answer.

A scale that varies smoothly across the sphere

The error is the simplest one real data have. A crystal that is longer in one direction than another absorbs X-rays more along its long axis. To lowest order the effect is a scale that depends on the direction of the reflection and varies over the sphere of directions like the second Legendre polynomial of the angle to that axis: too strong near the axis and opposite it, too weak in the band around it.

The shape of an uncorrected absorption error. Every reflection to two ångström with a non-negative third index, placed by its direction in a stereographic view of the upper hemisphere, and marked by how its measured amplitude is scaled: too strong near the fixed direction a and opposite it, too weak in the band around it, with the mark's size following the second Legendre polynomial of the angle to a. The direction is general, not along any axis of the cell, and the error is sized so that its root-mean-square effect is 3% of the mean amplitude — the same size as the random error it is set beside.
Fig. 1 Every reflection to two ångström with a non-negative third index, placed by its direction in a stereographic view of the upper hemisphere, and marked by whether the smooth error makes it too strong or too weak, with the mark’s size following the size of the effect. The error’s direction is general, not along any axis of the cell, and its root-mean-square effect is 3% of the mean amplitude.

The direction chosen is (1,2,3)(1, 2, 3) in the cubic cell’s axes, along no symmetry element, so the error breaks the cell’s symmetry exactly as a real crystal’s shape does. Its size is fixed by one condition: the root-mean-square change it makes to the amplitudes is three per cent of the mean amplitude. That is the size the random error had. So every comparison below sets two errors of the same size against each other, and differs only in whether the error is random or smooth.

Everything else is unchanged from the earlier refinements. The same eighteen pairs are used, drawn from three site symmetries, with the same reflection list of four hundred and sixty-two reflections to two ångström. The refinements are the same, with positions on each model’s own loci, two-component displacement parameters, occupancies fixed and four starts for the wrong model. The decision rule is also the same: a pair is decided when the likelihood ratio passes 10.83, the 0.1 per cent point of χ2\chi^2 on one degree of freedom. It is decided for the right model if the statistic is positive and for the wrong one if it is negative. In every case the error estimate used is the random one, three per cent of the mean, because it is the only estimate a refinement has. A smooth error the experimenter does not know about is, by definition, not in the weights.

At three per cent, the smooth error hardly matters

The first comparison is the one the earlier essay’s conclusion predicts should go badly. The data carry the smooth error and no random error at all.

A smooth error decides like a random one of the same size. Each of the eighteen pairs at sixty-four atoms, placed by its likelihood-ratio statistic with 3% random error across and with a 3% smooth scale error instead up, both computed with the same σ and on a signed logarithmic scale. The shaded bands are the undecided region, within 10.83 of zero. The points lie close to the dashed diagonal: the error enters both models' residuals equally and cancels from their difference, apart from its correlation with the models' own difference, which is small for an error that varies smoothly across reciprocal space.
Fig. 2 Each of the eighteen pairs at sixty-four atoms, placed by its likelihood-ratio statistic with 3% random error across and with the 3% smooth error instead up, both with the same σ and on a signed logarithmic scale. The shaded bands are the undecided region. The points lie close to the diagonal: the two kinds of error, at the same size, decide the same pairs.

The points lie along the diagonal. At sixty-four atoms the smooth error alone decides fourteen pairs for the right model, leaves four undecided, and decides none wrongly. The random error of the same size decides fifteen and leaves three. The median statistic is 112 under the smooth error and 92 under the random one. Where one error’s statistic is large, so is the other’s.

The reason is visible in the algebra of the statistic. After refinement, the right model’s residual on each reflection is whatever error it cannot absorb, ere_r. The wrong model’s residual is that same error plus the part of the two models’ difference it could not absorb, er+Dre_r + D_r. The likelihood ratio is the difference of their sums of squares,

∑r(er+Dr)2−∑rer2=∑rDr2+2∑rDr er,\sum_r (e_r + D_r)^2 - \sum_r e_r^2 = \sum_r D_r^2 + 2 \sum_r D_r\, e_r ,

divided by the error variance. The error’s own square cancels. What survives is the models’ difference and a cross term: the error multiplied by the models’ difference, reflection by reflection. The models’ difference changes sign from one reflection to the next, because it is made of partial atoms at different positions, while a smooth error hardly changes between neighbouring reflections. Their products therefore mostly cancel, and the cross term stays small compared with the first sum. A random error does the same, for a different reason: its values are unrelated to the models’ difference. At a given size the two cross terms are of the same order, and that is why the points sit on the diagonal.

With both errors present, the smooth one on top of the random one, the result at sixty-four atoms is fourteen pairs decided for the right model, three undecided and one decided wrongly. That one pair, a model keeping one subgroup of 4/mmm against another, has a statistic of −28.2-28.2. It is the first wrong decision any of these refinements has produced, and it arrives at a size where the earlier essay recorded none.

Where the two models’ difference falls through the error

The earlier essay’s argument was a comparison of sizes, and the sizes can be drawn.

The models' difference falls through the error's floor. For the eighteen pairs, the median residual between the two models' ideal averaged structures, which falls as one over the number of atoms from 9.7% at sixteen to 0.6% at two hundred and fifty-six, against the median residual the right model is left with when the only error is the smooth scale surface, which stays at 2.0% whatever the molecule. The two cross between sixty-four and a hundred and twenty-eight atoms.
Fig. 3 For the eighteen pairs, the median residual between the two models’ ideal averaged structures, falling as one over the number of atoms, against the median residual the right model keeps when the only error is the smooth surface, which stays at 2.0% for every molecule. The two cross between sixty-four and a hundred and twenty-eight atoms.

The models’ difference falls as one over the size of the molecule. It is 9.7 per cent at sixteen atoms, 2.4 per cent at sixty-four, 1.2 per cent at a hundred and twenty-eight and 0.6 per cent at two hundred and fifty-six. The smooth error leaves the right model a residual of 2.0 per cent whatever the molecule, because no structural parameter can follow a scale surface. The two lines cross between sixty-four and a hundred and twenty-eight atoms. The earlier essay expected the data to stop choosing at about that crossing, the same argument that had first put the boundary near sixty atoms.

The crossing is not where the decisions stop, and the cancellation is the reason. The statistic does not compare the models’ difference with the error. It compares the models’ difference with the error’s correlation with that difference, and for a smooth error that correlation is small. At a hundred and twenty-eight atoms the models are four tenths as far apart as the error is large, and with the random error at three per cent the smooth error still leaves twelve of the eighteen pairs decided correctly. That is one fewer than random error alone. A comparison of sizes predicts a collapse where the measurement shows a nudge.

The same crystal, counted better

The difference between the two kinds of error shows up when the experiment is improved in the way counting statistics improve: more exposure, more redundant measurements, a better detector. Every reflection, several times over showed that averaging redundant measurements reduces the random error and cannot touch an error common to all of them. So the next comparison holds the smooth error at three per cent and lowers the random error from three per cent to one, then to three tenths of one.

Better counting makes an uncorrected smooth error choose the wrong model. Eighteen pairs of disorder models, each decided by the likelihood ratio at the 0.1 per cent level, at sixty-four and a hundred and twenty-eight atoms, as the random error falls from 3% to 1% to 0.3% of the mean amplitude. In each group, the first bar has random error only, the second adds a smooth scale error of 3% across reciprocal space, and the third adds the same error with a degree-two scale surface refined alongside the structure. With random error only, better counting decides more pairs correctly. With the uncorrected smooth error it stops helping and the number decided for the wrong model grows, to three at a hundred and twenty-eight atoms. With the surface refined, the result is the random-only one in every group.
Fig. 4 The eighteen pairs at sixty-four and a hundred and twenty-eight atoms, as the random error falls from 3% to 1% to 0.3% of the mean amplitude. In each group of three bars: random error only; the smooth error added; the smooth error added and a scale surface refined with the structure. The lowest part of each bar is decided for the right model, the middle undecided and the top decided for the wrong one, with the count of wrong decisions above.

The picture at the head of this essay is that experiment. With random error only, better counting does what it should. At a hundred and twenty-eight atoms the right decisions go from thirteen to sixteen to seventeen as the error falls from three per cent to one to three tenths, and the only pair left undecided is the one no data can decide. With the smooth error present, the first step helps a little, from twelve to fourteen right decisions, and after that nothing more. Meanwhile the wrong decisions go from one to two to three. At sixty-four atoms the pattern is the same one size smaller: the right decisions stop at fifteen and the wrong ones rise to two.

Counting harder through an uncorrected smooth error makes the refinement confidently wrong. The algebra says why. Lowering the random error removes its part of the cross term, but the smooth error’s part stays, because it is the same in every measurement. The statistic divides the whole difference by the random variance, and that variance is what counting improves. For a pair whose smooth cross term outweighs the models’ difference, the difference is negative. It does not change as the counting improves, and dividing it by an ever smaller variance drives it past −10.83-10.83. The pair is decided for the wrong model at the 0.1 per cent level, and the confidence reported is a statement about the counting, not about the model.

What perfect counting would conclude

The limit of that process can be computed directly. With the random error removed entirely, the sign of the difference between the two refined residuals is the verdict any small enough random error would eventually reach.

What perfect counting would conclude through a smooth error. For each molecule size from sixteen to two hundred and fifty-six atoms, how many of the eighteen pairs have the right model fitting the data better than the wrong one when the data carry the smooth 3% scale error and no random error at all. That sign is the verdict the likelihood ratio approaches as the counting error shrinks to nothing. All eighteen are right at sixteen and thirty-two atoms, seventeen at sixty-four, fifteen at a hundred and twenty-eight and fourteen at two hundred and fifty-six.
Fig. 5 For each molecule size from sixteen to two hundred and fifty-six atoms, how many of the eighteen pairs have the right model fitting the smooth-error data better than the wrong one, with no random error at all. All eighteen at sixteen and thirty-two atoms, seventeen at sixty-four, fifteen at a hundred and twenty-eight and fourteen at two hundred and fifty-six.

At sixteen and thirty-two atoms all eighteen signs point the right way. At sixty-four one does not, at a hundred and twenty-eight three do not, and at two hundred and fifty-six four do not. One of the four is the pair that no data decide, whose two models fit equally well because the fitted model’s loci contain the true one’s. With no random error its sign is decided by nothing structural. The other three are genuine: pairs whose difference, at that size, happens to correlate with the scale surface more than it differs.

This is the measurement the earlier essay’s conclusion wanted, and it puts the boundary in the right place. The molecule size at which systematic error takes over is where perfect counting would begin to reach wrong verdicts. For a three per cent smooth error that is between thirty-two and sixty-four atoms, and by two hundred and fifty-six it is one pair in six. That is a much gentler boundary than a comparison of sizes predicts, and a much worse failure than “the data stop choosing”. Before it, the data choose correctly. After it, they choose with confidence, and sometimes wrongly.

Refining the error away

The correction crystallographers apply is itself a refinement. The scale is allowed to vary over the sphere of directions as a sum of spherical harmonics, and the coefficients are refined alongside the structure. That is the form the empirical absorption corrections in routine use take, a method Robert Blessing introduced in 1995. Here the scale is given the five real harmonics of degree two, which is exactly the family the error belongs to.

An empirical correction recovers the error it can represent. For one pair at sixteen atoms with the smooth error and no random error, the five coefficients of the degree-two scale surface refined alongside the right model's structure, set beside the true error surface projected on the same five harmonics by least squares. They agree to better than one part in ten billion, and the right model's residual returns to zero: a correction whose basis spans the error removes it exactly.
Fig. 6 For one pair at sixteen atoms with the smooth error and no random error, the five coefficients of the degree-two scale surface refined with the right model, beside the error surface projected on the same five harmonics. They agree to better than one part in ten billion, and the right model’s residual returns to zero.

The refinement recovers the surface: five coefficients, each matching the true one to eleven digits, and a residual for the right model that returns to zero on data with no random error. In the counting experiment the corrected refinements give the random-only answer in every one of the six groups: the same number of right decisions, the same undecided pair, no wrong decision anywhere. Five extra parameters are a small price for a data set of four hundred and sixty reflections, and they remove the whole effect.

That result depends on the correction’s basis spanning the error, which it does here by construction. A real crystal’s absorption surface is smooth but not purely of degree two, and extinction, which weakens the strongest reflections, follows the structure rather than the direction, so no expansion over directions represents it. Each is a separate correction with its own model. The argument above says what is at stake when one is left out: the sign of the statistic for the pairs whose difference correlates with what was left out. And it says that improving the counting cannot compensate.

Why this is not the earlier essay’s conclusion

The earlier conclusion was that the boundary is “set by the systematic error of the experiment and not by its counting statistics”. The measurement agrees that systematic error sets a limit, and disagrees about almost everything else. At the counting of a routine experiment, three per cent, the smooth error hardly matters: the decisions are set by the counting. The systematic limit appears only as the counting improves, and it appears as wrong verdicts rather than missing ones. So the practical advice reverses. A crystallographer who finds that two disorder models are undecided should not expect a longer exposure to settle the question. The longer exposure will settle it, but correctly only if every systematic error correlated with the models’ difference has been modelled. Otherwise it will settle it with growing confidence in whichever model that error favours.

The earlier essay has been corrected in place to say so, with a link here. Its other finding stands: the size of the partial atoms’ displacement parameters is a more robust signal than the residual, because a wrong model with nowhere to put an atom smears it whatever the error is.

What these refinements cannot show

A single error shape, at a single size, in a single general direction, does not survey systematic error. Other shapes correlate differently with the models’ differences, and the specific pairs that go wrong would change with the direction of the absorption axis. What does not change is the mechanism: the cancellation of the error’s square, the cross term, and the division by a variance that counting shrinks and the smooth error does not. The counts here are one sample of how large that effect is. They are not a table to look numbers up in.

The reflection list is complete to two ångström and noise-free apart from the errors put there deliberately. A real refinement weights reflections by their estimated errors, includes hydrogen atoms and a solvent model, and applies corrections that interact with each other. Any of these can absorb part of a wrong model or part of a systematic error. The figures show medians and counts over eighteen pairs. They do not show which pairs go wrong or why those pairs are the susceptible ones. That is the correlation between a particular disorder difference and a particular error surface, and it is not drawn here.

What the refinements have to refuse

What the systematic-error refinements must satisfy. Four tests, each able to fail: the random-error refinement must be unchanged by the new options, a smooth error alone must leave the right model a non-zero residual, and refining the degree-two surface must return it to zero. One claim is refused: that a smooth error the size of the random one can only blur a decision, when at sixty-four atoms it decides one pair for the wrong model.
Fig. 7 Four tests, each able to fail: the random-error refinement unchanged by the new options; a smooth error alone leaving the right model a residual; the refined degree-two surface returning it to zero. One claim refused: that a smooth error of the random error’s size can only blur a decision.

The refused claim is the natural reading of the earlier conclusion: that an error of the random error’s size, only smooth instead of random, can at worst make the data stop choosing. One pair at sixty-four atoms is decided for the wrong model with a statistic of −28.2-28.2, so the error’s effect is not limited to undecided pairs.

The first test guards the comparison itself. The refinement options added here, a systematic error in the data and a scale surface in the model, leave the random-error refinement exactly as it was, down to the statistic 196.3 on the pair the earlier essay reported. Without that test a change in the refinement itself could masquerade as an effect of the error.

Still open: which pairs, and which errors

The measurement says how many pairs go wrong and not which. The susceptible pairs are those whose difference, reflection by reflection, correlates with the error surface. That correlation can be computed from the two models’ ideal structures before any refinement, as an inner product of their difference with each harmonic of the correction’s basis. Whether it predicts which pairs go wrong, and so lets a crystallographer know in advance which corrections a given pair of disorder models is sensitive to, is the question the counts above point to and do not answer.

The other open direction is the error that no expansion over directions can model. Extinction weakens the strongest reflections by an amount that depends on their strength. It therefore correlates with the structure itself, and it is the systematic error most likely to correlate with a structural difference. Its effect on these decisions, and whether the usual one-parameter extinction correction removes it the way five harmonics removed the absorption, has not been computed.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Average structureDisorderDisplacement parameterLeast squaresOccupancyResolutionSite symmetryStructure factor