Unequal atoms break the equality, and very unequal ones mend it
Assumes The relation that is an equality, The average that knows the atoms and not where they are and A twin hides in the statistics.
The relation that is an equality found one relation among structure factors that is not a probability. For a cell of equal, resolved atoms, the square of the density has its peaks in the same places as the density, so each structure factor is exactly proportional to the convolution of all the structure factors with themselves, with a factor that depends on the atom’s shape and nothing else. The essay ended on a prediction it did not test: with two kinds of atom the square would no longer be proportional to the density, the logarithm of Sayre’s ratio would no longer be a straight line, the scatter about the line would measure the spread of atomic numbers — and, if so, the scatter would be an estimate of how heterogeneous a cell is, computable from amplitudes alone.
This essay runs that test. The first part of the prediction holds exactly, and in a form sharper than the prediction stated. The second part fails twice. The scatter does not measure the spread of atomic numbers: as one species gets heavier the spread grows without limit, while the scatter rises, peaks and falls back. And it cannot be computed from amplitudes, because Sayre’s ratio is built from phases. What the amplitudes can say about composition turns out to be a different number, which counts atoms rather than comparing them — and which, in a small cell with a heavy atom, reads exactly like a twin.
Why the square stops being proportional
The identity rests on one step. A density made of equal Gaussian atoms, squared, is a sum of narrower Gaussians at the same positions, plus cross terms that vanish when the atoms do not overlap. With atoms of different weights the squared density is still a sum of narrower Gaussians at the same positions, but the weight of each is now rather than . Its transform is therefore not proportional to the density’s. Writing for the phase factor of atom at reflection , the ratio Sayre’s identity computes becomes
The convolution on the left is the transform of the squared density, by the convolution theorem, so the ratio compares the density’s transform with its square’s, reflection by reflection. For resolved atoms both transforms are sums over the same atomic positions with the same phase factors; they differ only in the shape of each peak, which absorbs, and in the weight each atom is given, which nothing absorbs.
When every is the same, is the constant and the identity is the one already computed. Otherwise changes from reflection to reflection, in size and in phase, because the two sums weight the same atoms differently and so point in different directions in the complex plane.
The contrast is not subtle. The equal-atom cell sits on its line with a scatter of six thousandths in the logarithm, which is the overlap of Gaussian tails and the truncation of the convolution at a finite window. The cell with two heavy atoms scatters by nearly four tenths: individual reflections land a factor of one and a half above or below the line. The slope survives, since is still there and still falls with resolution; what fails is the claim that is the whole story.
It is worth being clear about what the scatter is not. It is not a failure of resolution, since the atoms are exactly as sharp and as far apart in the right-hand panel as in the left. It is not truncation of the convolution, since the window is the same. And it does not grow with resolution in the way overlap errors do: the points at high and low index scatter alike. Something about the cell’s contents, independent of how well they are resolved, has broken the proportionality, and the next figure says exactly what.
The scatter is the composition, exactly
The derivation above predicts more than that the line scatters. It predicts which way each reflection departs from it: by , a quantity computed from the atoms with no convolution at all. That is a claim a computation can check reflection by reflection.
The agreement is the sharpest part of the result. Sixty reflections, each of which required a convolution over some two thousand products of complex structure factors, lie on a line against a quantity that took one sum over eight atoms apiece. Nothing in Sayre’s ratio is left unexplained once the two weightings are written down. It also settles what kind of failure this is. Unequal atoms do not degrade the identity in some diffuse way that makes it unreliable everywhere. They replace a constant by a definite function of the structure, and the function is the ratio of two structure factors of the same atoms with different weights.
That has a consequence for the standard repair. The relation that is an equality noted that practical direct methods replace the amplitudes by normalised ones, dividing out the scattering of the cell’s contents shell by shell, and called that an approximation rather than a repair. The computation here says why it cannot be a repair. Normalising divides every atom’s contribution at a given resolution by the same number, so it changes nothing about the relative weights of the atoms, and the scatter is made of nothing but relative weights: in the density against in its square. A normalised structure factor of a two-species cell carries the same mismatch as the raw one, because the mismatch lives in the ratio of two sums over the same atoms and no common factor touches a ratio. What normalising does do is put reflections at different resolutions on one scale, which is what the statistical relations need, and that is a different job.
What sets the size of the scatter
The prediction said the scatter would measure the spread of atomic numbers. That turns out to be the wrong quantity, and the right one comes from asking how closely the two sums in follow each other as runs over reflections. For atoms at random positions their correlation over reflections is
which is one when every weight is equal, since then the two sums are proportional, and less than one otherwise. When is close to one the numerator and denominator of move together and barely changes; as falls they move independently and wanders. So the size of the scatter is governed by , the part of one sum the other does not explain.
Two properties of matter. It does not depend on the number of atoms: doubling a cell while keeping its composition doubles every sum and leaves the ratio unchanged. And it is not a measure of spread. Hold the fraction of heavy atoms at a quarter and make them heavier. The spread of weights — their standard deviation over their mean — grows steadily, from zero at equal weights to 0.74 at four times and 1.37 at sixteen. But first falls and then rises again, because once the heavy atoms are heavy enough the light ones stop contributing to either sum, and a cell dominated by two equal heavy atoms is, for this purpose, a cell of two equal atoms.
For a quarter of the atoms at weight and the rest at one, the correlation has a closed form,
which is one at , tends to one again as grows without limit, and dips in between to a minimum of 0.944. The dip is shallow as numbers go, but is what matters and it runs from nought to eleven hundredths and back.
The measured scatter follows the lower curve’s shape and not the spread’s. It climbs from the overlap floor as soon as the heavy atoms are half again as heavy as the rest, stays high between twice and six times, and falls by half or more at sixteen times, in both cell sizes. The lower curve’s maximum sits near a weight of 2.7 for a quarter of the atoms heavy; it can be found in closed form, and it is where the identity is worst for that composition. A cell of carbon with a quarter of its atoms replaced by sulphur, a weight ratio near 2.7, sits almost exactly there. A cell of carbon with a quarter replaced by lead is much further up the right-hand slope, and much kinder to the identity than the spread of its atomic numbers suggests.
That is the first way the prediction fails, and it is instructive rather than merely wrong. Sayre’s identity is broken by atoms that are comparable, not by atoms that are different. A structure of one very heavy species among light ones is close to a structure of that species alone, and the identity holds for it nearly as well as for equal atoms: the light atoms become a small perturbation on a density the heavy ones dominate. The heavy-atom method of phasing relies on the same fact from the other side: when the heavy atoms dominate the scattering, their positions alone fix most of every phase, and the light atoms are a correction.
The fraction matters as much as the weight, which the same formula makes easy to see. A single sulphur atom among twenty carbons gives a of 0.891, worse than the quarter-sulphur cell’s 0.944, because one atom of intermediate weight is a large perturbation that dominates nothing. A single lead atom among twenty carbons gives 0.958, and a quarter of lead gives 0.993. And the light atoms of an organic molecule on their own — six carbons, a nitrogen and an oxygen, weights six, seven and eight — give 0.993, close enough to one that the identity holds nearly as well as for equal atoms. That is one reason direct methods succeeded first on structures of light atoms: to Sayre’s identity, carbon, nitrogen and oxygen are very nearly the same atom.
The phases go wrong in the same place
Sayre’s identity is useful, when it is useful at all, because it relates phases: the ratio’s phase is nought for equal atoms, so the phase of each structure factor is fixed by the convolution of the others. With unequal atoms has a phase of its own, and that phase is the error a phase predicted from the identity would carry.
A root-mean-square phase error of half a radian, at twice the weight, is about thirty degrees on every reflection before any other source of error is counted. That is a large systematic error in a subject where the tangent formula refines phases towards a fixed point of relations like this one: the fixed point it finds is the fixed point of the wrong relation. Most of the phase relations that do not move with the origin descend from Sayre’s identity with terms thrown away, and they inherit this error, in a diluted form, whenever the atoms in a cell are of comparable but unequal weight.
What the amplitudes can and cannot say
The prediction’s second half was that the scatter could be read from amplitudes alone. It cannot. Sayre’s ratio divides a complex structure factor by a convolution of complex structure factors, and every term of that convolution carries two phases. Drop the phases and the convolution is not defined; keep only amplitudes and there is no ratio to scatter. The scatter is a measurement of the cell that can be made once the phases are known, which is to say once the structure is solved.
The amplitudes do see composition, through a different door. The average that knows the atoms showed that the mean intensity in a shell of reflections is the sum of the squared scattering factors, with every positional term averaged away. The next moment up, the mean of the squared intensity, keeps one term more: for a cell without a centre of symmetry, the fourth moment of the normalised amplitudes is
and that is a function of the atomic weights with no phase in it, measurable from intensities alone.
But it is not . The correction term falls as one over the number of atoms, so the fourth moment measures how few effective scatterers a cell has, which depends on its size as much as on its composition. The same composition at eight atoms and at thirty-two atoms gives the same , and the same Sayre scatter, and fourth moments of 1.57 and 1.89. The amplitudes count the atoms; Sayre’s scatter compares them. Neither can stand in for the other, and with two unknowns in a two-species composition — the fraction and the weight — one number could not have fixed both in any case.
For a centrosymmetric cell the untwinned value is three rather than two, which is what the centre test reads; the formula above is the acentric one, and every cell computed here is acentric.
The eight-atom value deserves a warning of its own. A twin hides in the statistics reads twinning from exactly this moment: an untwinned acentric crystal gives two, a perfect twin gives one and a half, and a twin fraction lowers it to . A cell of eight atoms with two of them six times heavier gives 1.57 untwinned, which that relation would read as a twin fraction of about 0.31. The estimator assumes the many-atom limit in which the correction term vanishes, and a small cell with a heavy atom is outside it. Knowing the cell contents, the term can be computed and subtracted; not knowing them, a heavy-atom structure and a twin are not distinguished by this moment.
What the computation has to refuse
Each claim above is computed from the atoms and the convolution directly, and each is tested against a case it would fail on if it were wrong.
The fourth test is the one that overturns the prediction, and it is set up so that it could have confirmed it instead. If the scatter measured the spread of atomic numbers, it would be larger at sixteen times the weight than at four, since the spread nearly doubles between them. It is smaller at both cell sizes, by roughly half. The last refusal carries the same kind of weight in the other direction: a composition read off the fourth moment without the cell size would put the eight-atom and thirty-two-atom cells at different heterogeneities, and they have the same one.
Still open: how much heterogeneity the phase relations tolerate
This computation measured how the exact identity degrades, and it did so with the phases in hand. The question that matters in practice is the downstream one: how much does a phase error of the size found here, systematic and concentrated where is lowest, slow or mislead the procedures that use the identity’s descendants? The solver that knows no symmetry finds structures of equal atoms from random phases by flipping the sign of weak density. Whether it finds a cell of carbon and sulphur as readily as a cell of carbon alone, and whether its success rate dips near a weight ratio of 2.7 and recovers beyond it as does, is a measurement this site’s solver could make on these same cells. The prediction, by analogy with everything above, is that it does, and the formula for says where to look: a single sulphur among twenty carbons should be harder for the solver than a quarter of sulphur, and a single lead atom easier than either, although the lead cell’s spread of atomic numbers is the largest of the three by far. The relation that can say no raises the same question for quartets, whose negative estimates depend on the cross terms being weak, and unequal atoms change which cross terms are weak.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- The phase problem direct methods · phase problem · structure factor
- A map of the atoms that break the law phase problem · structure factor
- One experiment gives the cosine, the other gives the sine phase problem · structure factor
- The reciprocal lattice phase problem · structure factor
- The zones that behave as if there were a centre phase problem · structure factor
- Two structures, one Patterson phase problem · structure factor
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
Direct methodsElectron densityNormalised structure factorPhase problemSecond momentStatisticsStructure factor