The phase problem
Every reflection a crystal produces has two numbers attached: how strong it is, and what its phase is. A detector counts photons, so it measures the first and not the second, and the phase of every reflection is gone before anything is recorded.
The three panels are the classical demonstration and the third is the one that makes the case. Discarding the amplitudes costs almost nothing; discarding the phases costs everything.
What is measured, and what is not
The structure factor of a reflection is a sum over the atoms in the cell:
which is a complex number: an amplitude and a phase . The intensity a detector records is proportional to , so the amplitude survives the measurement and the phase does not.
The loss is exactly half the information, in a literal sense — one real number recovered out of the two the complex value carries — and it cannot be recovered from more measurements of the same kind. Recording the same reflection a thousand times gives the same amplitude a thousand times over and never gives the phase.
Turning the measured amplitudes back into a structure requires the Fourier synthesis
and the sum needs . Every structure ever solved has come from supplying it by some means other than measuring it.
What the three panels do
The figure runs the synthesis three times on a computed structure, so the true phases are known and can be withheld deliberately.
True amplitudes, true phases. The synthesis reproduces the structure: the density peaks land on the atoms, and the check is exactly that — the highest points of the map are located and compared against the atom positions, with agreement required for every one.
True amplitudes, scrambled phases. Each phase is replaced by one that depends only on the reflection’s indices, so the substitution is deterministic and the figure is identical on every build. The map that results has peaks, because a sum of cosines always has peaks, and they are not at the atoms. The assertion here is a failure requirement: the count of peaks landing on atoms must come out below the number of atoms, and if it did not the demonstration would be worthless.
Unit amplitudes, true phases. Every is set to one and the phases are kept. The map is noisier — sharp features and a rougher background — and the peaks are still on the atoms, which the check confirms at the same tolerance as the first panel.
That third result is not obvious and it is the whole point. The amplitudes describe how strongly each spatial frequency is present; the phases describe where each one sits. Structure is a matter of alignment, and alignment lives entirely in the phases.
The same fact in other places
The phase problem is a special case of something general, and the general version is worth naming because it makes the result less surprising and more useful.
The same asymmetry appears in image processing. Take two photographs, Fourier-transform both, swap their phases while keeping their amplitudes, and transform back: each result looks like the picture whose phases were used. Edges, outlines and recognisable objects survive the phase and not the amplitude, for exactly the reason the third panel above works.
It appears in acoustics, where the relative phases of a sound’s harmonics have a much weaker effect on timbre than their amplitudes — the one common case where the asymmetry runs the other way, because the ear discards phase deliberately.
And it appears in astronomy, where interferometry recovers the amplitude of the visibility function easily and its phase only with difficulty, and where the techniques for coping have names — closure phase, self-calibration — that a crystallographer would find familiar.
How structures get solved anyway
Four families of method, in the order they were invented, and each supplies phases from something other than the measurement.
The Patterson function. Fourier-transform the intensities directly — no phases needed, since is what is measured — and what comes back is not the structure but the map of all interatomic vectors. For a structure with one heavy atom among light ones, the heavy-atom vectors dominate and its position can be read off, which gives approximate phases for everything. Patterson published this in 1934 and it carried the field for thirty years.
Isomorphous replacement. Measure the crystal, then measure it again with a heavy atom attached at a known site, and the difference in amplitudes constrains the phases. This is how the first protein structures were solved — myoglobin and haemoglobin, by Kendrew and Perutz — and Perutz spent more than twenty years getting it to work.
Anomalous scattering. Near an absorption edge an atom scatters with a phase shift of its own, which breaks the symmetry between a reflection and its opposite. The difference between the two is small and measurable, and it constrains the phase. Most protein structures today are solved this way, with selenium substituted for sulphur to supply the effect.
Direct methods. The phases are not arbitrary: the density must be positive everywhere, and it must be concentrated at atoms rather than spread out. Those constraints impose statistical relationships among the phases of related reflections — if two reflections have strong amplitudes, the phase of a third related one is probably close to their sum — and enough such relationships determine the whole set. Hauptman and Karle worked this out in the 1950s, the community ignored it for a decade, and they shared the Nobel Prize in Chemistry for it in 1985.
Direct methods are the interesting case for the argument of this site. They work because the answer is known to have properties the data do not enforce, and adding those properties recovers information the measurement lost. Nothing about the intensities says the density is positive; the chemistry does.
Reading the maps
The three panels reward a closer look, because the way each one fails is as informative as whether it fails.
The correct map has compact peaks on a flat background. Each peak is a single atom, its height is roughly the atom’s scattering power, and the space between peaks is close to zero. That is what an interpretable electron-density map looks like, and recognising one is the skill a crystallographer is trained in.
The scrambled map has peaks of comparable height in the wrong places, and — this is the part worth noticing — it does not look like noise. A sum of cosines with arbitrary phases produces a perfectly plausible-looking arrangement of blobs. Somebody handed the scrambled map with no other information would set about interpreting it, and would produce a structure.
The unit-amplitude map has peaks at the atoms with a rougher background around them. Flattening the amplitudes sharpens the peaks and adds ripples, because setting every to one is equivalent to weighting the high-resolution terms far more heavily than they deserve. It is a recognisable procedure — sharpening — and it is applied deliberately in practice, at a milder strength, to make peaks easier to pick out.
The middle panel is the one that should worry a reader most. A wrong answer to this problem is not obviously wrong, which is the same hazard a pattern figure with the wrong caption presents, in a field where the stakes are structures rather than illustrations.
Where the exactness stops
The figure’s claims are measurements with stated criteria, and the criteria matter more here than in most places on this site.
A synthesis is computed on a finite grid, so “the peak is on the atom” means “the highest cell of the map is within one and a half grid steps of the atom position”. That tolerance is stated in the code and printed in the assertion, and a coarser grid would loosen it.
The sum is over a finite set of reflections — everything inside a window in reciprocal space — which is the computational counterpart of an experiment’s resolution limit. Truncating the sum produces ripples around each peak, and at low resolution the ripples merge and the peaks broaden until individual atoms are no longer separable. Real structure determination lives with exactly this, and the resolution quoted with a published structure is the radius of that window.
And the scramble is a demonstration rather than a proof. Showing that one wrong assignment of phases destroys the map is not showing that no wrong assignment could accidentally reproduce it — that is a statement about all possible phase sets, and no figure establishes it. What the figure establishes is that the phases are not redundant, which is the claim it is captioned with.
The Patterson map, which is what the data alone give
There is exactly one map that can be computed from a diffraction measurement with no extra information, and it is worth knowing what it contains because it marks the boundary of what the experiment supplies.
Fourier-transform the intensities rather than the amplitudes — that is, use with all phases set to zero — and the result is the Patterson function. It is not the structure. It is the map of every interatomic vector: a peak at the position corresponding to each pair of atoms, with height proportional to the product of their scattering powers.
Two consequences follow immediately. A structure with atoms gives a Patterson map with off-origin peaks, so the map is far more crowded than the structure and rapidly becomes uninterpretable as grows. And the map is always centrosymmetric, whether or not the structure is, because for every vector between two atoms there is the opposite vector between the same two.
What makes it usable is contrast. A single heavy atom among many light ones contributes peaks proportional to the square of its scattering power, which dominate everything else, so its position can be extracted from the crowd. That is the entire basis of heavy-atom phasing, and it is why the technique needs an atom conspicuously heavier than the rest rather than merely a labelled one.
The honest summary is that the data alone give the autocorrelation of the structure, and an autocorrelation is a structure with its phases removed — which is the same statement as the one at the top of this page, arrived at from the other side.
The surprising part
The phase problem sounds like a technical obstacle and it is closer to a structural feature of what measurement is.
Here is the connection worth carrying. An intensity is a squared modulus, and squaring destroys sign information in exactly the way that makes ambiguous. A diffraction experiment measures for each reflection independently, and what it cannot see is the relationship between reflections — because a phase is meaningful only relative to another phase, and each intensity is measured alone.
So the missing information is relational rather than local. Each individual reflection is measured perfectly well; what is lost is how they fit together. And every method for solving the problem works by importing a relationship from somewhere else: a heavy atom in a known place, an anomalous scatterer, a positivity constraint, a known homologous structure.
That is the same shape as the argument this site makes about patterns and their groups. A symmetry is a relationship between points, so no amount of examining points individually finds it, and the detector works by testing candidate relationships exhaustively. Diffraction has the harder version of the same problem, since the relationships it needs are not drawn from a finite list and cannot be enumerated.
What symmetry gives back
Symmetry helps, and it helps in a way worth stating because it connects this essay to the rest of the site.
A centrosymmetric structure — one with an inversion centre — has real structure factors, so every phase is or . The continuous phase problem collapses to a binary one, and a structure with a thousand reflections has possibilities rather than a continuum. That is still large and it is enormously better, and it is why centrosymmetric structures were solved decades before non-centrosymmetric ones of comparable size.
Every symmetry element imposes relationships among phases of related reflections, so the higher the symmetry, the fewer independent phases there are to find. And the systematic absences that identify the space group also reduce the count of reflections whose phases are needed.
Symmetry is therefore not merely a description of the answer; it is part of the machinery for finding it. A crystallographer determines the space group first, before attempting the structure, and does so because it is the cheapest information available.
Who solved what
Patterson’s function dates from 1934 and was the first method that worked without guessing. Perutz’s heavy-atom work on haemoglobin ran from 1937 to the late 1950s, and its length is a fair measure of the problem’s difficulty.
Hauptman and Karle’s direct methods, from 1953, were the change of regime. Their central paper was received with scepticism — the objection was that phases could not possibly be derivable from intensities, which is true as stated and false once positivity is added — and the methods now solve essentially every small-molecule structure automatically, in seconds, with no heavy atom and no second crystal.
The most recent shift is computational rather than mathematical. Structure prediction from sequence has become accurate enough that a predicted model can supply starting phases for a measured protein, which is molecular replacement with a computed rather than an experimental starting point. The phase problem is not solved; it is increasingly answered from outside the experiment, which is what every method in this essay has done.
Where the ladder goes next
The geometry the phases are attached to is the reciprocal lattice, and the information that survives without any phasing at all is the systematic absences.
The further loss an experiment can suffer is what a powder pattern loses, where orientation goes as well.
The construction the phases are used with is the dual lattice, and the symmetry information that reduces the problem comes from the classification.
What the pictures here cannot show. The three maps on this page are syntheses from a structure that was known in advance, which is precisely the situation a real experiment is never in. What a crystallographer has is the middle panel’s information — amplitudes and no phases — with no way to tell which of the three the answer resembles. The figure demonstrates what the phases are worth; it cannot demonstrate the difficulty of not having them.