The two most anomalous wavelengths are the wrong pair
Assumes One experiment gives the cosine, the other gives the sine, How much of it is the other hand and The law that hides handedness.
One experiment gives the cosine, the other gives the sine solved the phase problem for one reflection by intersecting two circles. An isomorphous replacement fixed the cosine of the phase, and an anomalous difference fixed its sine. Each alone left a two-fold ambiguity, and the two ambiguities sat at right angles, so together they left none. That essay named multi-wavelength anomalous diffraction, MAD, as the experiment that gets both from one crystal. Near an absorption edge the same atoms supply the isomorphous-like difference, by changing their real scattering with wavelength, and the anomalous one, through the imaginary part. It computed nothing about it.
What a MAD experiment has to decide is which wavelengths to use. The absorption edge of the anomalous atom, selenium in most macromolecular work, offers a handful of distinguished energies: the peak of the imaginary part , the inflection where the real part dips furthest, and remote energies well away from the edge on either side. The natural guess is to measure where the anomalous effects are largest, at the peak and the inflection. That guess is wrong, and this essay shows why with the equations the experiment actually solves. Of all the pairs of wavelengths that use the edge at all, the peak and the inflection together recover the phase worst.
The equations one reflection supplies
The anomalous atoms scatter , where is their ordinary scattering and and change sharply near the edge. Write for the whole structure’s ordinary scattering, for the anomalous atoms’ ordinary part, and for the phase difference between them. Then each measured intensity, for a reflection and its Friedel mate at wavelength , has the form Jerome Karle wrote down in 1980:
with , and . The unknowns are , and . Knowing and the anomalous atoms’ positions, which a map of the atoms that break the law finds from the anomalous differences themselves, gives the phase of .
The structure of the equations settles most of what follows, and it can be read off directly. Adding a Friedel pair removes the sine. Subtracting it isolates the sine, with the coefficient , proportional to . So the sine of the phase is carried by alone, at whichever wavelength has the most of it. The cosine has no such separate channel. It appears in the sums, alongside , and at a single wavelength it cannot be told apart from them. Two wavelengths give two sums, and their difference removes and leaves together with a small term in . The cosine is carried by the change of between the two wavelengths, the dispersive difference. A pair of wavelengths whose values are close determines the cosine badly, however large is at either.
One wavelength is not enough, and counting the equations shows exactly what is missing. A single wavelength gives one sum and one difference: two equations for three unknowns. The difference fixes . The sum fixes one combination of , and the cosine term, and nothing in one wavelength separates the cosine from the other two. In practice the amplitudes are taken from elsewhere, from the mean of the pair and from the anomalous atoms’ positions. The sine then fixes the phase difference only up to the choice between and , which share every sine. That is the two-fold ambiguity of single-wavelength anomalous diffraction, SAD. The earlier essay noted that in practice it is resolved by density modification, choosing between the two maps by how flat the solvent looks. A second wavelength resolves it inside the data instead, by giving the cosine a channel of its own. How well it does so is set by how different the second wavelength’s is.
With two wavelengths there are four intensities and three unknowns, so the problem is over-determined and non-linear. Here it is solved for the three physical unknowns by least squares from eight starting phases, keeping the best fit. With no measurement error the solution is exact: sixty reflections at the peak and a remote wavelength, all recovered to within radians.
An edge, and f′ from f″
The wavelength choice needs and across the edge, and the two are not independent. Causality ties the real part of a response to its imaginary part through the Kramers–Kronig relation, so a model that assumed both would be two assumptions where physics allows one. Here is modelled, and is computed from it.
The model is a stated shape, not measured data. rises from half an electron below the edge to about 3.8 above it, with a white-line peak of 5.4 electrons three electronvolts above. Selenium’s real edge has roughly that form. follows from the transform: a dip to electrons at the edge, the inflection, rising to about electrons at 150 electronvolts on either side. The transform is a principal-value integral done numerically. Its constant, which the relation leaves free, is fixed by setting to electron far below the edge.
The transform has to be checked on something with a known answer before its output is trusted. A Lorentzian absorption line has a dispersion partner in closed form, and the numerical transform reproduces it at six energies across the line to within seven thousandths of an electron. The four distinguished wavelengths then have definite pairs of values. The peak has and . The inflection has and . The high remote has and , and the low remote and .
Every pair, phased
With the edge in hand, every pair of the four wavelengths can be tried on the same synthetic data. Each experiment has two hundred reflections with Wilson-distributed amplitudes, the anomalous atoms carrying two per cent of the scattering power, and a one per cent error on every intensity. Each pair recovers the phase difference reflection by reflection, and the median error measures the pair.
The picture at the head of this essay ranks them, and the ranking follows the right-hand column, the change of , more closely than anything about . The inflection paired with either remote is best, at seven degrees: those pairs have the largest spread of , over five electrons. The peak paired with either remote is next, at eight degrees, with a spread of about four. The peak and the inflection together, the two wavelengths with the largest anomalous effects, give eighteen degrees, more than twice as bad. Their values are only electrons apart, because they are three electronvolts apart on the same edge. The two remotes, whose values differ by a third of an electron, are useless at fifty-five degrees. Adding the high remote to peak and inflection helps, to degrees, which is why careful experiments measure three wavelengths when the crystal survives it.
Where the error goes
The ranking can be explained term by term. The phase is recovered as an angle from its cosine and sine, and each can be scored separately against the truth.
For the peak and the inflection the cosine is the weak term. Its median error is , against for the inflection with the high remote, while the sine errors are and . The sine is well determined whenever one wavelength has plenty of , and the peak has the most. The cosine needs the sums at two wavelengths to differ, and they differ in proportion to the change of . For the two remotes the cosine error reaches , which is nearly no information. The sine degrades too, even though the high remote has almost four electrons of , because the least-squares solution couples the terms. An undetermined cosine lets and drift, and the sine is read against them.
So the right choice is a matter of complementarity, the same principle the earlier essay found between isomorphous replacement and anomalous scattering, where two constraints at right angles leave no ambiguity. Here it is inside one experiment. One wavelength should have a large for the sine, and the pair should span a large change of for the cosine. The inflection has the most negative and a substantial . A remote wavelength has an close to its far value. Together they supply both. The peak has the most , but its is close to the inflection’s, and pairing the two spends both measurements on the same edge feature.
The solver is not the reason
A ranking of pairs measured with one solver could be a ranking of how well that solver copes with each pair rather than of the pairs themselves. An iterative least-squares fit started from eight phases might simply be worse at finding the answer for some combinations of coefficients. The check is independent of any solver. For each reflection, the Fisher information of the four measured intensities about the three unknowns can be computed from the equations and the error model alone. Its inverse gives the smallest variance any unbiased estimate of the phase could have: the Cramér–Rao bound.
The bound and the solver agree. For the inflection with the high remote, the median error the bound allows is 7.6 degrees and the solver achieves 7.2. For the peak with the inflection they are 17.4 and 18.4, and for the peak with the high remote 7.4 and 8.2. The solver is within a few tenths of a degree of the best possible on the pairs that are any good, so the ranking belongs to the wavelengths. The bound is not meaningful for the two remotes, since it is a statement about small errors and their errors are not small. There the solver’s fifty-five degrees are already most of the way to random, which is ninety.
The agreement also says where the eighteen degrees come from. The bound is set by the inverse of the Fisher information, and a pair whose two sums differ only slightly gives a nearly singular information matrix in the direction of the cosine. That is the dispersive difference again, now as the smallest eigenvalue of a matrix rather than a difference of two numbers.
How much selenium
The last variable the experimenter does not choose, but the chemist sometimes can, is how much anomalous scattering there is. Here the anomalous atoms carry two per cent of the scattering power. That is the order of one selenomethionine for every hundred and fifty or so residues, the common case. Varying it with the wavelengths fixed at the inflection and the high remote shows how the error depends on it.
At half a per cent the median error is 14 degrees, at one per cent 10, at two per cent 7.2, at four 4.8 and at eight 3.5. Quadrupling the anomalous scattering halves the error, so the phase error goes as one over the square root of the anomalous fraction. That means one over the anomalous amplitude, because the fraction is a ratio of intensities. This is the scaling one would expect from the equations, in which every phase-dependent term is proportional to . It puts the choice of wavelengths in proportion. The factor of two and a half between the peak-and-inflection pair and the best pair is what quadrupling the selenium content would buy, or reducing the counting error by the same factor. Choosing the second wavelength well costs nothing and is worth as much as either.
Counting harder does not change the choice
A reasonable objection is that all of this is about noise: with good enough data any pair would do. The objection is half right. With perfect data every pair except the degenerate ones recovers the phase exactly, as the noiseless test shows. The question is how each pair’s error scales with the counting error.
For the choices that include the edge the lines are nearly straight on logarithmic axes, with slope close to one. Halving the error on each intensity halves the phase error, whatever the pair, so the ratio between two pairs stays nearly constant as the data improve. At 0.2 per cent the inflection and high remote give 1.3 degrees and the peak and high remote 1.8. At 4 per cent they give 26 and 29. The choice of wavelengths is a constant factor that counting cannot remove. Every reflection, several times over showed that redundancy buys a smaller random error and nothing else, and here that smaller error is multiplied by a factor fixed when the wavelengths were chosen.
That also explains why the recommendation has been stable since MAD became routine. Wayne Hendrickson and his colleagues developed it around 1990, with selenium introduced into proteins as selenomethionine. The usual advice for two wavelengths pairs the peak or the inflection with a remote energy, and never the two edge features together. The reason is not a rule of thumb. It is the structure of the equations: the sine lives at one wavelength and the cosine between two.
What the model leaves out
The synthetic experiment is idealised in the ways that matter least to the comparison and most to absolute numbers. Real intensities carry systematic errors, including absorption, radiation damage between the two wavelengths and imperfect scaling between datasets. The dispersive difference is especially sensitive to these, because it compares two datasets rather than two mates within one. Damage accumulated between the first and second wavelength adds an error that looks exactly like a change of . That is the practical argument for measuring the wavelengths interleaved, and it could shift the balance towards single-wavelength anomalous diffraction for fragile crystals. That shift is outside this model.
The edge is also a model. Real selenium edges vary with the chemical environment. The white line can be taller or absent, and the inflection sharper or broader. A sharper edge separates the peak and inflection less in energy and more in , which would narrow the gap between the peak-and-inflection pair and the others. The ranking here is for this edge. What carries over to every edge is the mechanism: the cosine follows the change of .
The phase difference is not yet a phase. Turning into the phase of needs the anomalous atoms’ positions, found from the anomalous or dispersive differences by the Patterson methods of a map of the atoms that break the law. Errors in those positions add to the errors here. The comparison between pairs assumes the positions are known exactly.
What the calculation has to refuse
The refused claim is the intuition this essay began with: that the wavelength with the largest anomalous signal is what matters, and the second wavelength hardly does. With the peak fixed, the second wavelength changes the median error from 8.2 degrees to 18.4, depending on whether it is a remote energy or the inflection. The test that makes the whole comparison trustworthy is the noiseless one. Without it, an error in the solver would be indistinguishable from an ill-conditioned pair.
Still open: three wavelengths, and where the third should go
Adding the high remote to the peak and the inflection brought the error to 5.7 degrees, the best of anything tried. That leaves an optimisation this essay has not run. Given three wavelengths, where on the edge should each go? The complementarity argument suggests that the third should extend the spread of , since is already well supplied by the peak, but a low remote does that almost as well as a high one and has little of its own. A search over all triples of energies across the edge, with a model of radiation damage that penalises the dataset measured last, would turn the qualitative advice into a number. The same calculation with four wavelengths would say whether the fourth is ever worth its dose.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- Better counting, and the wrong model least squares · resolution · structure factor
- The zones that behave as if there were a centre friedel law · phase problem · structure factor
- What refinement takes back from a wrong model least squares · resolution · structure factor
- Where the experiment runs out anomalous scattering · friedel law · structure factor
- A merohedral twin moves no spot at all friedel law · structure factor
- How many reflections there are friedel law · resolution
The objects this essay names
Each one links to every other essay that touches it.
Anomalous scatteringFriedel lawLeast squaresPhase problemResolutionStructure factor