The phase problem
Assumes The reciprocal lattice and Systematic absences.
Every reflection a crystal produces has two numbers attached: how strong it is, and what its phase is. A detector counts photons, so it measures the first and not the second, and the phase of every reflection is gone before anything is recorded.
The three panels are the classical demonstration and the third is the one that makes the case. Discarding the amplitudes costs almost nothing; discarding the phases costs everything.
What is measured, and what is not
The structure factor of a reflection is a sum over the atoms in the cell:
which is a complex number: an amplitude and a phase . The intensity a detector records is proportional to , so the amplitude survives the measurement and the phase does not.
The loss is exactly half the information, in a literal sense — one real number recovered out of the two the complex value carries — and it cannot be recovered from more measurements of the same kind. Recording the same reflection a thousand times gives the same amplitude a thousand times over and never gives the phase.
Turning the measured amplitudes back into a structure requires the Fourier synthesis
and the sum needs . Every structure ever solved has come from supplying it by some means other than measuring it.
What the three panels do
The figure runs the synthesis three times on a computed structure, so the true phases are known and can be withheld deliberately.
True amplitudes, true phases. The synthesis reproduces the structure: the density peaks land on the atoms, and the check is exactly that — the highest points of the map are located and compared against the atom positions, with agreement required for every one.
True amplitudes, scrambled phases. Each phase is replaced by one that depends only on the reflection’s indices, so the substitution is deterministic and the picture is the same every time it is drawn. The map that results has peaks, because a sum of cosines always has peaks, and they are not at the atoms. The assertion here is a failure requirement: the count of peaks landing on atoms must come out below the number of atoms, and if it did not the demonstration would be worthless.
Unit amplitudes, true phases. Every is set to one and the phases are kept. The map is noisier — sharp features and a rougher background — and the peaks are still on the atoms, which the check confirms at the same tolerance as the first panel.
That third result is not obvious and it is the whole point. The amplitudes describe how strongly each spatial frequency is present; the phases describe where each one sits. Structure is a matter of alignment, and alignment lives entirely in the phases.
The same fact in other places
The phase problem is a special case of something general, and the general version is worth naming because it makes the result less surprising and more useful.
The same asymmetry appears in image processing. Take two photographs, Fourier-transform both, swap their phases while keeping their amplitudes, and transform back: each result looks like the picture whose phases were used. Edges, outlines and recognisable objects survive the phase and not the amplitude, for exactly the reason the third panel above works.
It appears in acoustics, where the relative phases of a sound’s harmonics have a much weaker effect on timbre than their amplitudes — the one common case where the asymmetry runs the other way, because the ear discards phase deliberately.
And it appears in astronomy, where interferometry recovers the amplitude of the visibility function easily and its phase only with difficulty, and where the techniques for coping have names — closure phase, self-calibration — that a crystallographer would find familiar.
How structures get solved anyway
Four families of method, in the order they were invented, and each supplies phases from something other than the measurement.
The Patterson function. Fourier-transform the intensities directly — no phases needed, since is what is measured — and what comes back is not the structure but the map of all interatomic vectors. For a structure with one heavy atom among light ones, the heavy-atom vectors dominate and its position can be read off, which gives approximate phases for everything. Patterson published this in 1934 and it carried the field for thirty years.
Isomorphous replacement. Measure the crystal, then measure it again with a heavy atom attached at a known site, and the difference in amplitudes constrains the phases. This is how the first protein structures were solved — myoglobin and haemoglobin, by Kendrew and Perutz — and Perutz spent more than twenty years getting it to work.
Anomalous scattering. Near an absorption edge an atom scatters with a phase shift of its own, which breaks the symmetry between a reflection and its opposite. The difference between the two is small and measurable, and it constrains the phase. Most protein structures today are solved this way, with selenium substituted for sulphur to supply the effect.
Direct methods. The phases are not arbitrary: the density must be positive everywhere, and it must be concentrated at atoms rather than spread out. Those constraints impose statistical relationships among the phases of related reflections — if two reflections have strong amplitudes, the phase of a third related one is probably close to their sum — and enough such relationships determine the whole set. Hauptman and Karle worked this out in the 1950s, the community ignored it for a decade, and they shared the Nobel Prize in Chemistry for it in 1985.
Direct methods are the interesting case for the argument of this site. They work because the answer is known to have properties the data do not enforce, and adding those properties recovers information the measurement lost. Nothing about the intensities says the density is positive; the chemistry does.
Reading the maps
The three panels reward a closer look, because the way each one fails is as informative as whether it fails.
The correct map has compact peaks on a flat background. Each peak is a single atom, its height is roughly the atom’s scattering power, and the space between peaks is close to zero. That is what an interpretable electron-density map looks like, and recognising one is the skill a crystallographer is trained in.
The scrambled map has peaks of comparable height in the wrong places, and — this is the part worth noticing — it does not look like noise. A sum of cosines with arbitrary phases produces a perfectly plausible-looking arrangement of blobs. Somebody handed the scrambled map with no other information would set about interpreting it, and would produce a structure.
The unit-amplitude map has peaks at the atoms with a rougher background around them. Flattening the amplitudes sharpens the peaks and adds ripples, because setting every to one is equivalent to weighting the high-resolution terms far more heavily than they deserve. It is a recognisable procedure — sharpening — and it is applied deliberately in practice, at a milder strength, to make peaks easier to pick out.
The middle panel is the one that should worry a reader most. A wrong answer to this problem is not obviously wrong, which is the same hazard a pattern figure with the wrong caption presents, in a field where the stakes are structures rather than illustrations.
Where the exactness stops
The figure’s claims are measurements with stated criteria, and the criteria matter more here than in most places on this site.
A synthesis is computed on a finite grid, so “the peak is on the atom” means “the highest cell of the map is within one and a half grid steps of the atom position”. That tolerance is stated in the code and printed in the assertion, and a coarser grid would loosen it.
The sum is over a finite set of reflections — everything inside a window in reciprocal space — which is the computational counterpart of an experiment’s resolution limit. Truncating the sum produces ripples around each peak, and at low resolution the ripples merge and the peaks broaden until individual atoms are no longer separable. Real structure determination lives with exactly this, and the resolution quoted with a published structure is the radius of that window.
And the scramble is a demonstration rather than a proof. Showing that one wrong assignment of phases destroys the map is not showing that no wrong assignment could accidentally reproduce it — that is a statement about all possible phase sets, and no figure establishes it. What the figure establishes is that the phases are not redundant, which is the claim it is captioned with.
The Patterson map, which is what the data alone give
There is exactly one map that can be computed from a diffraction measurement with no extra information, and it is worth knowing what it contains because it marks the boundary of what the experiment supplies.
Fourier-transform the intensities rather than the amplitudes — that is, use with all phases set to zero — and the result is the Patterson function. It is not the structure. It is the map of every interatomic vector: a peak at the position corresponding to each pair of atoms, with height proportional to the product of their scattering powers.
Two consequences follow immediately. A structure with atoms gives a Patterson map with off-origin peaks, so the map is far more crowded than the structure and rapidly becomes uninterpretable as grows. And the map is always centrosymmetric, whether or not the structure is, because for every vector between two atoms there is the opposite vector between the same two.
What makes it usable is contrast. A single heavy atom among many light ones contributes peaks proportional to the square of its scattering power, which dominate everything else, so its position can be extracted from the crowd. That is the entire basis of heavy-atom phasing, and it is why the technique needs an atom conspicuously heavier than the rest rather than merely a labelled one.
The honest summary is that the data alone give the autocorrelation of the structure, and an autocorrelation is a structure with its phases removed — which is the same statement as the one at the top of this page, arrived at from the other side.
The surprising part
The phase problem sounds like a technical obstacle and it is closer to a structural feature of what measurement is.
Here is the connection worth carrying. An intensity is a squared modulus, and squaring destroys sign information in exactly the way that makes ambiguous. A diffraction experiment measures for each reflection independently, and what it cannot see is the relationship between reflections — because a phase is meaningful only relative to another phase, and each intensity is measured alone.
So the missing information is relational rather than local. Each individual reflection is measured perfectly well; what is lost is how they fit together. And every method for solving the problem works by importing a relationship from somewhere else: a heavy atom in a known place, an anomalous scatterer, a positivity constraint, a known homologous structure.
That is the same shape as the argument this site makes about patterns and their groups. A symmetry is a relationship between points, so no amount of examining points individually finds it, and the detector works by testing candidate relationships exhaustively. Diffraction has the harder version of the same problem, since the relationships it needs are not drawn from a finite list and cannot be enumerated.
What symmetry gives back
Symmetry helps, and it helps in a way worth stating because it connects this essay to the rest of the site.
A centrosymmetric structure — one with an inversion centre — has real structure factors, so every phase is or . The continuous phase problem collapses to a binary one, and a structure with a thousand reflections has possibilities rather than a continuum. That is still large and it is enormously better, and it is why centrosymmetric structures were solved decades before non-centrosymmetric ones of comparable size.
Every symmetry element imposes relationships among phases of related reflections, so the higher the symmetry, the fewer independent phases there are to find. And the systematic absences that identify the space group also reduce the count of reflections whose phases are needed.
Symmetry is therefore not merely a description of the answer; it is part of the machinery for finding it. A crystallographer determines the space group first, before attempting the structure, and does so because it is the cheapest information available.
The phases are not equal within an orbit, and that is the point where a reader can lose the thread. An operation shifts a phase by a known amount, so one phase in an orbit determines the rest exactly; what the group buys is not equal phases but a smaller set of unknowns. The higher the symmetry, the fewer independent phases there are to find, and the count above is the count of both.
The extreme case of that is a centre of symmetry, and it is worth seeing on its own because it changes the kind of problem rather than its size.
That is the whole of what a centre buys. A thousand reflections give two-to-the-thousand sign assignments rather than a continuum of angles — still an enormous number, and enormously better, and it is why centrosymmetric structures were solved decades before non-centrosymmetric ones of comparable size.
Who solved what
Patterson’s function dates from 1934 and was the first method that worked without guessing. Perutz’s heavy-atom work on haemoglobin ran from 1937 to the late 1950s, and its length is a fair measure of the problem’s difficulty.
Hauptman and Karle’s direct methods, from 1953, were the change of regime. Their central paper was received with scepticism — the objection was that phases could not possibly be derivable from intensities, which is true as stated and false once positivity is added — and the methods now solve essentially every small-molecule structure automatically, in seconds, with no heavy atom and no second crystal.
The most recent shift is computational rather than mathematical. Structure prediction from sequence has become accurate enough that a predicted model can supply starting phases for a measured protein, which is molecular replacement with a computed rather than an experimental starting point. The phase problem is not solved; it is increasingly answered from outside the experiment, which is what every method in this essay has done.
Where the ladder goes next
The geometry the phases are attached to is the reciprocal lattice, and the information that survives without any phasing at all is the systematic absences.
The further loss an experiment can suffer is what a powder pattern loses, where orientation goes as well.
The construction the phases are used with is the dual lattice, and the symmetry information that reduces the problem comes from the classification.
What the pictures here cannot show. The three maps on this page are syntheses from a structure that was known in advance, which is precisely the situation a real experiment is never in. What a crystallographer has is the middle panel’s information — amplitudes and no phases — with no way to tell which of the three the answer resembles. The figure demonstrates what the phases are worth; it cannot demonstrate the difficulty of not having them.
Why the problem is soluble at all
Half the information is gone, permanently, and structures are solved every day. The reconciliation is a counting argument, and it is worth doing because it says which structures are easy and which are not.
Count the unknowns. A structure with atoms in the asymmetric unit is described by coordinates, plus a displacement parameter or six for each — call it four to nine numbers per atom. A small organic molecule with thirty independent atoms needs a few hundred numbers.
Count the measurements. The number of unique reflections inside a given resolution is fixed by the cell volume, and for that same crystal it runs to several thousand. The ratio is comfortably better than ten measurements per parameter.
So the problem is heavily overdetermined even after losing half of it. Losing the phases does not leave too little data; it leaves plenty of data in a form from which the answer cannot be read directly. That is a different difficulty, and it is why the methods listed above work by supplying phases from elsewhere and then refining against everything.
Which explains where the difficulty actually bites. A protein has thousands of independent atoms and diffracts to poorer resolution, so its ratio of measurements to parameters is far worse — sometimes below one — and the phases cannot simply be found by trying. That is why protein crystallography developed heavy-atom and anomalous methods into an industry while small-molecule crystallography solves structures automatically.
And it explains why the constraints matter so much. Electron density is positive everywhere and concentrated at atoms, and neither fact appears in the measurement. Both are extra information, worth as much as data, and the direct methods are built on them.
Sampling more finely than the crystal allows
There is a way to make the problem soluble outright, and it requires giving up the crystal.
A crystal samples the transform. Its diffraction exists only at reciprocal-lattice points, so the continuous transform of one cell’s contents is measured at a grid of points and nowhere between them. The spacing of that grid is set by the cell.
A single non-periodic object has a continuous transform, which can be measured as finely as the detector allows. Measure it on a grid twice as fine as the object’s own size would dictate and the number of measured intensities is more than twice the number of unknowns — at which point the phases are determined by the data, not merely constrained by it.
Iterative algorithms then find them. Guess the phases, transform back, impose what is known about the object — that it is zero outside a known region, that its density is positive — transform forward, keep the measured amplitudes, and repeat. The procedure converges for objects of ordinary complexity, and it is how images are formed at X-ray free-electron lasers from single particles that were never crystallised.
The trade is exposure. One molecule scatters some times more weakly than a crystal of them, which is the reason crystallography was invented and the reason the alternative had to wait for sources bright enough to destroy the sample in the time it takes to measure it.
What this makes readable
Essays that name this one as a prerequisite.
- The map that needs no phases
- The solver that knows no symmetry
- As sharp as the sphere is wide
- One experiment gives the cosine, the other gives the sine
- The formula that has the answer already
- The law that hides handedness
- The zones that behave as if there were a centre
- Three phases that do not move when the origin does
- Two structures, one Patterson
- Whether there is a centre is a statistic
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- The relation that can say no direct methods · phase problem · structure factor
- The relation that is an equality direct methods · phase problem · structure factor
- Solving from the vector set the patterson function · phase problem
- Symmetry does not rescue a Patterson the patterson function · structure factor
What links here
The 8 essays that link to this one and share the most of its objects, of 26 that link here.
- A map of the atoms that break the law
- One experiment gives the cosine, the other gives the sine
- The solver that knows no symmetry
- Three phases that do not move when the origin does
- Two structures, one Patterson
- When the atoms are not all the same
- Where the pairs come from
- The map that needs no phases
The objects this essay names
Each one links to every other essay that touches it.
Direct methodsFourier synthesisThe Patterson functionPhase problemStructure factor