The heavy atom the flipping solver cannot see past
Assumes Unequal atoms break the equality, and very unequal ones mend it, The solver that knows no symmetry and The relation that is an equality.
Unequal atoms break the equality measured how Sayre’s identity comes apart when a cell holds two kinds of atom. The ratio the identity makes constant for equal atoms scatters, reflection by reflection, and the size of the scatter follows with
a correlation that is one for equal atoms, falls to a minimum near a weight ratio of 2.7 for a quarter of heavy atoms, and climbs back towards one as the heavy atoms come to dominate. That essay ended with a prediction about the solver that exploits the same sparseness of density: charge flipping, which finds structures from random phases by reversing the sign of weak density. A cell of carbon with a quarter of sulphur, , should be harder for it than carbon alone. A single sulphur among nineteen carbons, , should be harder still. A single lead, , should be easier, and the success rate should dip near a weight of 2.7 and recover.
This essay runs the solver on those cells. Every part of the prediction that concerns the solver fails. The sulphur cells are solved more often than carbon alone, not less. The lead cell is never solved, at any threshold, from any start. Success does not dip where dips; it collapses where recovers. The quantity that governs Sayre’s identity is not the one that governs the solver, and the one that does is a property of the map rather than of the relation.
Two things the test needed first
The first attempt at the measurement failed for a reason that had nothing to do with solving, and it is worth describing because it would have produced a wrong answer that looked like a finding. The solver’s success is judged by whether the twenty strongest peaks of its final map sit on the twenty atoms, allowing the origin and the hand to be chosen freely. With point atoms, the map made from the true phases of the lead cell fails that test.
The reason is truncation. A point atom scatters equally at every angle, and a map built from reflections out to a finite resolution is the structure convolved with the transform of that cutoff, which rings with side lobes a fifth as high as the central peak. A lead atom weighs eighty-two, so its side lobes are about as high as a carbon’s peak, and they compete with the carbons for the twenty strongest places. A solver judged against that map would be marked wrong for finding the true structure. So every cell here is given atoms whose scattering falls off with angle, with the fall-off about a tenth at the data’s edge, as a real atom’s electrons make it. With that done, the true map of every cell passes.
The second requirement is statistical. A success rate over random starts on one arrangement of atoms is a fact about that arrangement as much as about its composition, and the first single-arrangement runs disagreed with each other by more than the effects being sought. Every rate reported is therefore pooled over five random arrangements of the same composition, each tried from eight random phase sets, at each of seven thresholds.
Sulphur helps, lead stops everything
The picture at the head of this essay is the result. The cell of twenty carbons is solved from 13 of 40 starts at its best threshold, one standard deviation of the map. The cell with a quarter sulphur is solved from 24 of 40, at a slightly lower threshold of 0.9. The cell with a single sulphur is solved from 22 of 40. The prediction ran the other way on both: it made the sulphur cells harder because their is lower. In fact they are easier, and the cell with the lowest of all, one sulphur at 0.891, is solved nearly twice as often as the cell whose is exactly one.
The lead cell is solved from none of the forty starts at any of the seven thresholds, and a further sweep at thresholds down to a twentieth of a standard deviation solved none either. Its is 0.960, closer to one than either sulphur cell. If governed the solver, lead would sit between carbon and the sulphur cells, and it sits alone at zero.
How sure a rate over forty starts is
A success rate is a count, and forty starts do not pin a proportion down finely. A true rate of a third gives counts out of forty with a standard deviation of three, and a true rate of three fifths gives one of about three as well. So 13 of 40 for carbon and 24 of 40 for a quarter sulphur differ by about two and a half standard deviations of their difference, a gap chance would produce about once in a hundred tries. Carbon against a single sulphur, 13 against 22, differs by about two, which chance would produce about once in twenty-five. The direction of both comparisons is secure at this size. How much easier the sulphur cells are is not: a factor of 1.7 or 1.8 at the best thresholds, known to perhaps a third of its size.
The lead cell’s zero is secure in a different way. Forty failures in forty starts at the best threshold, and none at six other thresholds either, leave a true rate below about one in thirteen with ordinary confidence. More to the point, the explanation below predicts a rate of nought, not merely a small one. The thresholds tried span the whole range at which the other cells succeed, and the lead cell’s carbons sit inside that range rather than above it.
Pooling over arrangements matters as much as the count. An arrangement with two atoms nearly on top of each other, or with several atoms close to one line of the grid, is harder for the solver whatever the composition, and a rate measured on one arrangement would rank the cells partly by that accident. The single-arrangement runs made on the way to this measurement did not agree in their ranking of the carbon and sulphur cells, which is why the rates reported are pooled over five. A larger pool would narrow the rates further; whether it would change the ranking of the two sulphur cells, which differ by two counts in forty, it cannot be said from these.
Where the threshold lands
The explanation is in what the threshold is measured against. Charge flipping reverses the sign of every point of the map below standard deviations, and the standard deviation is the map’s own, the only scale an algorithm working on unscaled amplitudes has. For the algorithm to work, the atoms’ peaks must stand clearly above the threshold and the noise between them must fall below it.
In the cell of carbon alone, the carbons’ peaks stand a median of five standard deviations above the mean over the five arrangements. With a quarter sulphur the carbons sit lower, a median of three, because the sulphurs now contribute most of the variance; with a single sulphur they sit at four and a quarter. That is still well clear of the working threshold of 0.9, and the sulphurs themselves stand far above it. In the lead cell the lead owns the variance almost entirely, and the carbons stand at a median of 1.5 standard deviations, most of them between one and two. The threshold that would flip the noise between atoms flips the carbons as well, and a threshold low enough to spare the carbons flips almost nothing. There is no window, and the solver finds the lead and nothing else. That is not a failure of phasing, since the lead’s phases are nearly the whole phase set. It is a failure to see anything past the lead.
This also explains why sulphur helps. A few moderately heavy atoms make the map sparser in the way the algorithm relies on: a small number of strong peaks and a large flat region, with the lighter atoms still clearly above the threshold. A plausible reading, not measured here, is that the flipping locks onto the heavy atoms first and their phases then carry the light atoms with them; what is measured is only that the rates rise. The effect Sayre’s ratio measures, the phase error a mixed cell introduces into the identity’s descendants, is real, but charge flipping does not use the identity. It uses the map, and what matters to it is the map’s dynamic range.
Through the minimum of ρ
The weight sweep that essay asked for makes the separation between the two quantities plain.
Five heavy atoms among fifteen light ones, pooled over four arrangements and six starts each. At a weight of one the cell is equal atoms and is solved from 10 of 24 starts. At 1.5 it is solved from 16, at 2 from 14 and at 2.7, where is at its minimum of 0.944, from 15. At 4 it falls to 11. At 6 and at 10, where has recovered to 0.973 and 0.988, it is solved from none. The solver’s success does not dip at the minimum of ρ; it survives it, and fails exactly where ρ recovers. Over the same sweep the light atoms’ peaks in the true map fall from 5.1 standard deviations at a weight of one to 3.8 at two, 2.3 at four, 1.6 at six and 0.9 at ten. The collapse comes when that number passes below about two, and the best threshold slides downwards ahead of it, from 1.0 to 0.9 to 0.8, until there is no threshold left.
So and the solver answer different questions. says how well the amplitudes of the density predict the amplitudes of its square, which is the relation direct methods based on triplets rely on. The light atoms’ height in standard deviations says whether a threshold can separate them from the noise, which is what a flipping solver relies on. The two move together only at first. A cell whose heavy atoms dominate is nearly a cell of equal heavy atoms, so for Sayre’s relation it is easy again. For a flipping solver it is the hardest cell of all, because everything that is not heavy has become noise.
The older method sees the carbons twice
A cell that one method cannot solve is not unsolvable, and the lead cell is exactly the case the oldest method of all was invented for.
The heavy-atom method takes the phases a single heavy atom would give, puts the measured amplitudes on them and makes a map. The heavy atom’s position comes from the map that needs no phases, where a heavy atom’s vectors stand out. For the lead cell this map places a peak on 18 of the 19 carbons. It also places one on all 19 of the carbons’ images inverted through the lead, while 19 positions shifted at random land on peaks only 6 times, which is what chance allows among forty peaks. The method finds the structure, and it finds its mirror image through the lead with equal confidence.
The reason is symmetry the heavy atom brings with it. A single atom is its own inversion through its own position, so the phases it gives are those of a centrosymmetric structure, and a map made with centrosymmetric phases is itself centrosymmetric about the heavy atom. The structure and its inverted image appear at equal height, and the phases cannot say which is the crystal. This is the false centre of symmetry that a single heavy atom imposes in a cell with no symmetry of its own, familiar in practice for as long as heavy-atom methods have been used. It is resolved by taking a few of the stronger light-atom peaks as correct, which breaks the tie, and recycling. The two methods thus fail in complementary ways on the lead cell. Charge flipping cannot see past the heavy atom at all. The heavy-atom method sees everything twice.
What this says about the prediction’s premise
The prediction was not careless. It reasoned that a cell whose Sayre ratio scatters is a cell whose phase relations are noisy, that every direct method descends from those relations, and that noisier relations make a harder problem. The first two steps are right. The third assumes that charge flipping uses the phase relations, and it does not, at least not directly. It uses a single property of the density, that most of the cell is nearly empty, and it enforces that property by reversing the sign of whatever is weak. Sayre’s identity is one consequence of a density being a sum of separated atoms; the flipping rule is a different consequence of the same fact, and there is no reason for the two to degrade at the same rate as the atoms become unequal.
The methods that do descend from the relations, the tangent formula and its successors in the formula that has the answer already, weight each triplet of reflections by a probability computed from the normalised amplitudes. Those probabilities do feel the composition. The normalisation divides by the scattering power of the whole cell, and a heavy atom changes both the normalised amplitudes and the reliability of every triplet. Whether a triplet-based solver dips at the minimum of , as the prediction expected of charge flipping, is the same measurement with a different solver, and it is the one that would test the premise directly.
What the census has to refuse
The first test is the control that saved the measurement, stated so that it cannot silently change. The last is the counterpart every solver needs: a map made from the measured amplitudes on random phases must fail the test the solutions pass, or the test would reward any map at all. Both hold, and the test is exactly as demanding as it was in the essay that introduced the solver.
Two conventions sit under every rate and should be named. The success test asks for the atoms among the strongest peaks, to within a grid step, with origin and hand free. A looser test, asking only for the heavy atoms, would credit the lead cell with success every time, and would be a test of a different thing. And the threshold is stated in standard deviations of the map, which is how the method was published and is used. A threshold stated against the light atoms’ expected peak height, if that could be known, would move the lead cell’s window back into existence. That is the idea behind variants of the method that rescale the map before flipping, and it is not tested here.
Still open: a threshold that knows the composition
The failure on the lead cell is a failure of one convention, the threshold measured in standard deviations of a map one atom dominates. Two repairs suggest themselves and neither has been run. One flips against a threshold scaled to the weakest atom the composition says is present, which the chemistry supplies before any phasing. The other removes the heavy atom’s own contribution from the map before each flip, once its position is known, so that the remaining variance belongs to the light atoms. The second is the heavy-atom method and charge flipping combined, and it would test whether the two complementary failures above cancel. Both measurements are the same pooled experiment as here, run on the lead cell with a changed flipping rule, and both would say whether the collapse beyond a weight of four is a property of the problem or of one way of solving it.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- The relation that can say no direct methods · normalised structure factor · phase problem · structure factor
- Three phases that do not move when the origin does direct methods · normalised structure factor · phase problem · structure factor
- A map of the atoms that break the law heavy atom · phase problem · structure factor
- The phase problem direct methods · phase problem · structure factor
- When the atoms are not all the same heavy atom method · phase problem · structure factor
- How many reflections it takes to know there is a centre heavy atom · normalised structure factor
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
Direct methodsHeavy atomHeavy atom methodNormalised structure factorPhase problemPseudosymmetryStructure factor