The molecule size that hides a disorder
Assumes The occupancy does not name the disorder, How many orientations a disorder needs and The symmetry of an average.
The occupancy does not name the disorder counts the disorder models at a site and finds that many of them share one occupancy — eleven at a single kind of tetragonal site, all at one quarter. It also says what separates them: which of the site’s operations the molecule keeps, which the averaged structure records and the occupancy does not. And it ends on the question a count of subgroups cannot answer. A structure refined with one model keeps its occupancies at a quarter by construction; a second model with the same occupancy and different special positions may fit nearly as well, and whether the two are separated in practice, at the precision a diffraction experiment reaches, is a question about data.
It turns out to have a number attached to it, and the number is a molecule size.
For a molecule of ten atoms, two models sharing an occupancy differ by about twenty per cent in their calculated amplitudes — several times any reasonable measurement error, so the data choose between them decisively. For a molecule of sixty, the same two models differ by about three per cent, which is the measurement error, and the data stop choosing. Seven pairs out of two hundred and thirty-eight differ by nothing whatever, at any size and any resolution.
What an averaged structure actually is
The average over the orientations has a form that makes the whole question tractable, and it is worth deriving rather than assuming. A molecule keeping a subgroup of its site’s symmetry takes orientations at occupancy , so the average is over coset representatives . Expanding each coset,
because every coset contributes identical copies of and . So the average is the -orbit of each of the molecule’s atoms, each image at occupancy one over the length of its own orbit — and the subgroup has vanished from the formula.
mmm, seen down the third axis, with circle area standing for occupancy. The faint atoms are the images of the molecule’s general-position atoms and are the same for every model at this site; the dark ones sit on loci the model keeps.has not really gone. It returns through where the atoms are. An atom in a general position of gives images at each, and gives the same images whichever subgroup the molecule keeps. An atom lying on a rotation axis or a mirror plane belonging to gives a shorter orbit at a correspondingly higher occupancy — and the loci available to lie on are exactly the loci of . That is where the model is written down in the average, and it is nowhere else. It is also why the question is about special positions rather than about molecules: the Wyckoff positions of a group are the list of orbit types, and the average can carry information about a model only by putting occupancy on one rather than another.
The consequence is the one the earlier argument states and this one measures: two models differ only in a handful of partial atoms near the site. Everything else about the two averaged structures is identical, atom for atom and occupancy for occupancy.
The comparison this sets up, and why it is the hard case
To measure the difference, each model is given the same molecule: a fixed number of atoms at one general position’s orbit, common to every model and identical between them, plus one atom’s worth of scattering distributed evenly over the loci the model keeps. Every model then has the same composition and the same total scattering, so the structure factor at zero scattering angle is exactly equal for all of them and no difference anywhere can be a difference of composition dressed up as a difference of structure.
A real reorientation would move the general atoms too, and would be much easier to see. Holding them fixed isolates the part of the difference that the group theory forces and nothing else, which is precisely the case the question is about: two models that put the bulk of the molecule in nearly the same place.
Structure factors are then computed on a cubic cell of side 12 Å with the site at the origin, unit scattering power for every atom and an isotropic displacement factor, out to 1 Å resolution — about three and a half thousand reflections. The cubic cell restricts the sites that can be used: a trigonal or hexagonal site’s operations are not isometries of a cube, and comparing them there would compare the wrong structures, so five of the nineteen site symmetries the census of sites turns up are excluded and named rather than quietly dropped. Nine of the remaining fourteen carry two models of one occupancy, and those nine are the survey.
One more thing is fixed by that setup and worth naming, because it is the reason the comparison has any power at all. Both models are built over the same site, so both have the same set of orbits available and the same reflection list. Nothing is being compared across cells or across space groups, where a residual would be contaminated by everything else that differs; the only thing that can move is which orbits carry the special occupancy. That is the narrowest possible comparison and therefore the most demanding one — any difference it finds is a difference the group theory put there.
Twenty per cent, which is a great deal
4/mmm, the closest and the farthest, each drawn as the loci its model keeps, with the residual between the two averaged structures beside it.The residual between two calculated sets of amplitudes is the ordinary crystallographic one, , and it is what a refinement would report if data generated by one model were fitted with the other. Across all 238 pairs the medians per site run from 8 to 24 per cent.
That is a very large number by the standards of the experiment. A small-molecule structure determination routinely reaches a residual of three to five per cent against its own data; a difference of twenty per cent between two candidate models is not a marginal call, it is two structures that no crystallographer would confuse. So the first answer to the question is reassuring: for a molecule of ten atoms, the models the subgroup census distinguishes are distinguished by the data, comfortably.
The spread within a site is wide — at 4/mmm the pairs run from nothing to 53 per cent — and the closest pairs at several sites sit uncomfortably near the noise even at ten atoms. Those are the pairs whose kept loci overlap heavily: two models that keep the same principal axis and differ only in a mirror plane, for instance, put most of their partial atoms in the same orbits and differ only in the rest.
The wide spread is itself the useful observation. A crystallographer facing eleven models at one occupancy is not facing eleven equally difficult decisions: most of them are separated by half the residual of a completely wrong structure, and a handful are separated by a few per cent. Knowing in advance which pairs are the close ones is a calculation on the site symmetry alone, made before any data exist — which is what counting the models gives and what the residuals here rank.
The difference is not hiding at high resolution
A plausible guess about a difference confined to a few partial atoms is that it lives at high resolution, where the data are weakest — that the models are separated only by the reflections a crystal is least willing to give.
The guess is wrong, and the reason is worth stating because it is the same reason twice. The residual is a ratio: the numerator falls as the scattering angle grows and so does the denominator, because both are built from the same atoms with the same displacement factor. So the residual is nearly flat, between 18 and 22 per cent in every shell from 12 Å to 1 Å.
What does fall is the significance. Noise in a diffraction experiment is not a fixed fraction of each amplitude; it is closer to a fixed amount set by counting statistics across the whole data set, and the high-resolution amplitudes are small against it. So the same relative difference clears three standard deviations in three reflections out of four at low resolution and in three out of five at high. The discrimination is a low-resolution measurement, which is the opposite of what a reader would guess and the opposite of what it would be if the models differed in fine detail rather than in occupancy on a few orbits.
That also explains something practical. A structure collected only to 2 Å — a protein data set, or a crystal that diffracts badly — has already got most of the information that separates the models, because the shells that carry it are the ones every experiment collects. Cutting the resolution costs the count of reflections steeply, since the number available grows as the cube of the limit, and costs the model discrimination hardly at all.
Where the data stop choosing
The residual is large because the partial atoms are a large share of a ten-atom molecule. They are a fixed budget of one atom’s worth however large the molecule is, so their share falls as with the number of atoms in general positions, and the residual falls with it.
The measured factor is 0.50 for each doubling, which is the law to two figures. For the pair drawn here the residual is 22 per cent at eight atoms, 12 per cent at sixteen, 6 per cent at thirty-two and 3 per cent at sixty-four — and the share of reflections separating the models by three standard deviations falls from three quarters to almost none across the same range.
So there is a molecule size at which the question changes character. Below about thirty atoms the models are separated by the data and the crystallographer’s choice between them is a measurement. Above about sixty they are not, and a structure reported with one of them has had the choice made by the refinement’s starting point, by chemical judgement, or by whichever model the software’s automatic treatment of a special position happened to build. All three are respectable ways to decide something; none of them is the data deciding.
It is worth being precise about what “the data stop choosing” means, because it is not that the structure is wrong. Both models fit; both give a chemically sensible arrangement; both reproduce the amplitudes to within the measurement error. What is lost is the claim that the reported model is the one the crystal has, and the report usually does not distinguish the two. The same distinction runs through the unknowns against the observations, where the question is whether there are enough measurements to determine the parameters at all; here there are thousands of measurements per parameter and the difficulty is that the parameters barely change the measurements.
The boundary moves with data quality in the obvious direction and moves with the number of special atoms in the less obvious one. A molecule that happens to put three atoms on its kept axis rather than one triples the budget and pushes the boundary out by the same factor, so a disorder model involving a heavy atom on a special position is separable at sizes where a model involving a hydrogen is not. That is the ordinary asymmetry of a diffraction experiment — what a structure factor weights is electrons and not atoms — arriving in a place where it decides a question of symmetry rather than of position.
Seven pairs no experiment reaches
Seven of the 238 pairs sit at a residual of exactly zero, and they are not a matter of precision. Their averaged structures are the same structure, so no resolution and no data quality separates them. Two mechanisms produce them, and both are instructive.
The site symmetry’s own operations may carry one model’s loci onto the other’s. At a cubic site, the four-fold axis along and the three two-fold axes along , and lie in the same orbit of the site group, because the site group permutes the three cell directions. So a model keeping 4/m and a model keeping mmm put their partial atoms on the same set of positions, and with the same composition on both they put the same occupancy there too. The models are genuinely different — the subgroups are not conjugate, and the molecules that realise them are differently shaped — and the average has no way to say so.
And a rotoinversion fixes nothing but the site itself. A four-fold rotoinversion carries a point on its own axis to the point opposite, so no atom can sit on it except at the centre — and the centre is a position every model has equally. A model keeping therefore leaves exactly the trace that a model keeping its square, the two-fold, leaves already. That is a clean statement about what an averaged structure can record: only the loci that atoms can occupy, and a rotoinversion has none of its own. It is the same fact that makes a rotoinversion nearly invisible in a table of site symmetries except through the rotation inside it, and it is the structural counterpart of the observation in what a molecule gives up to sit in a crystal that a site offers symmetry a molecule may decline to use.
What the comparison assumes, and what it cannot say
Three limits deserve to be stated rather than left for a reader to find.
The special scattering is spread evenly over a model’s loci, and that is a convention. A molecule keeping three two-fold axes need not have an atom on each, and if the occupancies on the three orbits are unequal the two models of an identical pair separate again. The seven pairs are therefore pairs a symmetric molecule cannot distinguish, which is the natural case and not the only one. What is not a convention is the equal total scattering: without it, a difference in composition would be counted as a difference in structure, and the comparison would flatter itself.
The noise model is crude. A single relative error applied to a mean amplitude is not what a data set looks like; real standard uncertainties vary reflection by reflection, rise at high resolution and rise for weak reflections. The direction of every conclusion here survives that, since the significance already falls with resolution, but the boundary in atoms is a scale rather than a threshold.
And the residual is not a refinement. Two models fitted to the same data would each adjust their free parameters, and the residual between the two fitted structures is smaller than the residual between the two ideal ones — a wrong model with adjustable positions and displacement parameters can absorb a good deal of the difference. That effect works in the same direction as everything above and makes the boundary tighter, not looser, which is why the numbers here are best read as the most favourable case.
Still open: what a refinement would actually do with the second model
The gap between the two ideal structures is computed. The gap that matters to a crystallographer is between the two refined structures — one model fitted freely to data generated by the other — and that is a least-squares problem rather than a group-theoretic one. It has a definite answer for each of the 238 pairs, it would be smaller than the numbers here by a factor nothing on this page predicts, and the factor is the whole of the difference between “the models differ” and “the models are distinguished”.
What can be said in advance is where to look. The parameters a wrong model has available to absorb the difference are the positions and displacement parameters of its own partial atoms, and those are few and heavily constrained by the site symmetry — a partial atom on a two-fold axis has one free coordinate rather than three. So the absorption should be poor, and the second question this leaves is whether the displacement parameters of the partial atoms are where a wrong model betrays itself: an atom refined on the wrong locus with a suspiciously large and anisotropic displacement is the standard diagnostic in practice, and whether it is a reliable one is a question the comparison above could be pointed at and has not been.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- Symmetry does not rescue a Patterson enumeration · orbit · special position · structure factor
- A form is an orbit, and whether it closes is an integer question multiplicity · orbit · special position
- Every reflection, several times over multiplicity · orbit · special position
- One part in however many, and why it is never quite that multiplicity · orbit · special position
- An orbit is what the invariants cannot tell apart orbit · special position
- Every colour count at once enumeration · orbit
The objects this essay names
Each one links to every other essay that touches it.
Average structureDisorderEnumerationMultiplicityOccupancyOrbitResolutionSite symmetrySpecial positionStructure factor