Near-symmetry, and the tolerance that is not here
Assumes The symmetry diffraction adds and What a symmetry actually is.
Everything else on this site rests on a property the subject is unusually lucky to have. A pattern’s symmetry operations are integer matrices in the lattice basis and its translations are rationals with small denominators, so “does this pattern have a mirror” is settled by comparing whole numbers. There is no residual, no threshold, and no judgement.
Measured structures do not arrive that way. They arrive as coordinates with uncertainties, refined against data, and the same question then has to be answered by comparing a distance against a number somebody picked.
The figure is the argument. The number of symmetry operations a structure has is not a property of the structure once the coordinates are measurements; it is a function of the tolerance, and the correct value is one step of a staircase with nothing in the data to mark it.
What was actually done
The construction is deliberately simple, so that nothing in the result can be blamed on the elaborateness of the method.
A p4m pattern is generated exactly, in the usual way: a motif, the group’s operations, and the round trip checking that the point set has precisely the eight operations it was built with. Then every atom is displaced in a fixed pseudo-random direction by up to a set fraction of a cell edge — one and a fifth per cent in the figure above, which is a plausible size for the difference between a real structure and its idealised symmetric description.
The displaced coordinates are then handed to a detector that accepts an operation when every atom’s image lands within the tolerance of some atom. That is the same search the exact detector runs, with one comparison changed from equality to nearness. Nothing else differs.
Three regimes, and only the middle one is right
Reading the staircase from left to right gives three answers to one question.
Below the displacement, the structure has no symmetry. At zero tolerance the count is one — the identity and nothing else. That is not an artefact: it is exactly true. Not one of the eight operations maps the displaced set onto itself, because every atom is in the wrong place by a little.
Across a broad band, it has eight. Once the tolerance exceeds roughly the displacement, all eight operations are accepted together and the count sits at the group’s own order across a wide range of thresholds. That plateau is the reason the practice works at all: the right answer is stable over a range, rather than being visible only at one magic value.
Above that, it has more than eight. Loosen far enough and operations start fitting that the pattern never had, and the count climbs. There is no upper regime in which it settles down again; every further loosening admits more.
So the answer as a function of tolerance is a staircase whose steps are 1, then 8, then a climb. The middle step is the group the pattern came from. Nothing in the coordinates says so. Knowing which step is right requires knowing what the displacement was, which is exactly what a real measurement does not tell.
A second structure, and the plateau moves
One staircase could be a special case, so here is another group at a different displacement.
The two staircases together say something the first alone could not. The plateau’s height is a property of the structure and its position is a property of the error. An analyst who knew the displacement could pick a threshold with confidence; an analyst given only coordinates has to infer the displacement from the same data that the threshold is meant to interpret, which is circular in exactly the way that makes this hard.
There is a practical rule hiding in the two pictures, and it is the rule real software uses: look for the plateau rather than for a value. Sweep the tolerance, and if the accepted count is stable across a wide band, that band’s value is worth believing. A count that changes with every step of the sweep is a count with nothing behind it. That converts the arbitrary choice into an observation, and it is the best available answer short of bringing in the data the coordinates were refined against.
The exact case, for comparison
It is worth putting the exact version beside all this, because the contrast is the point of the essay and the exact version is easy to take for granted.
The whole of that certainty comes from working in the lattice basis. In Cartesian coordinates a threefold rotation involves √3, every comparison becomes a floating-point comparison, and the same staircase appears immediately — not because the pattern has changed but because the description has. The exactness is a property of the coordinates chosen, and it is available whenever a pattern is specified rather than measured.
How large is a real displacement
The numbers used above are not arbitrary, and it is worth saying where they sit relative to practice.
A well-refined small-molecule structure has coordinate uncertainties of a few thousandths of an ångström, against cell edges of five to twenty ångströms — so a fraction of a cell edge in the fourth decimal place. At that precision the plateau begins almost immediately and the exact answer and the tolerant one agree.
The interesting cases are not there. They are structures where the deviation from the higher-symmetry description is real and small: a phase transition just below its transition temperature, a structure with a slightly ordered arrangement of atoms that the higher symmetry would average over, a molecule whose own symmetry is broken by its packing. The displacement is then a physical fact worth reporting rather than an error to be absorbed, and choosing a tolerance large enough to recover the higher symmetry means choosing to describe the structure as the thing it nearly is.
That decision is a scientific one and not a numerical one. The staircase is where it becomes visible.
The part that was not expected
The count does not rise monotonically. At four of the twenty-five steps in the figure above it falls as the tolerance is loosened.
The reason is worth following, because it is the point at which the analogy between the exact question and the tolerant one breaks down completely. With a tolerance, two operations that differ by a little are both accepted — and they are the same operation as far as the question is concerned, so they have to be merged before anything is counted. The merging threshold is the same tolerance. So loosening it does two things at once: it admits more operations, and it collapses more of the admitted ones together.
Which effect wins depends on the structure and on where the threshold sits, and near a plateau the merging can win. The count is therefore not monotone in the tolerance, which means it is not even usefully wrong in a predictable direction. An experimenter who found eleven operations at one threshold and nine at a looser one would be right to be alarmed and would have found nothing but this.
This was not the expected result. The figure was written to show a staircase and it showed a staircase with dips in it, which is a better figure and a worse advertisement for tolerance-based symmetry detection.
What “the same operation” even means now
The exact detector has no merging step, because two operations are equal or they are not. Introducing a tolerance destroys more than the sharpness of the answer; it destroys the algebra underneath it.
The accepted set need not be a group. If two operations are each accepted because their errors are within the threshold, their composition can have twice the error and fail. A set of operations that is not closed under composition is not a group, and every theorem in this subject is about groups. The site’s exact machinery cannot produce such a set — closure is how a group is generated in the first place — and the tolerant version produces one routinely.
A near-symmetry has no orbit. The orbit of a point is the set of its images under the group, and it is well defined because applying two operations in either order gives the same answer. Under a tolerance, applying operations in different orders gives answers that differ by accumulated error, so “the orbit” is a cloud whose size depends on how many operations were composed.
This is the deepest difference between what this site does and what a laboratory does, and it is not a matter of precision. The exact question has an answer that is a group; the tolerant question has an answer that is a list.
What crystallography actually does
The practice has developed answers to all of this, and they are not thresholds chosen better.
Refine in the higher symmetry and see if it holds. Rather than asking whether coordinates have a symmetry, a crystallographer imposes the candidate symmetry, refines the structure again with fewer free parameters, and asks whether the fit got worse by more than the loss of parameters explains. That converts a threshold question into a statistical test, and the answer comes with a stated confidence rather than a stated tolerance.
Look for pseudosymmetry as a warning sign. A structure that is very nearly of a higher-symmetry group is the standard way to get a published structure wrong: the wrong group fits acceptably, the coordinates come out slightly displaced from the truth, and nothing in the residual announces the error. Detecting missed symmetry has dedicated software, and corrections to published structures are a regular feature of the literature. That is the same hazard ornament runs into with symmetric motifs, arriving through a completely different route.
State the tolerance. The convention that helps most is the cheapest: report the threshold used. A symmetry claim with an unstated tolerance cannot be checked or reproduced, which is the methodological problem that makes surveys of decorative patterns disagree as well.
What the software actually does
Symmetry detection in measured coordinates is a solved practical problem, in the sense that programs exist and are used daily, and describing what they do makes the difference between this essay’s toy and the real thing clear.
The standard approach does not start with a tolerance sweep. It starts from a hypothesis: given a structure described in some group, is there a larger group it could be described in? The candidates are not arbitrary — they are the supergroups of the current group, which is a short list read off the classification — so the search is over a handful of possibilities rather than over all operations.
For each candidate, the structure is transformed into that group’s setting, the atoms are averaged over the orbits the larger group would require, and the displacement each atom has to move is computed. The output is a maximum displacement per candidate: this structure would be p4m if every atom moved by at most so much. That number is reported, and a person decides.
Three things follow from that design and all three are improvements on the sweep.
The comparison is per atom rather than global. One badly placed atom is visible as one large displacement, where a global count of accepted operations hides it entirely.
The candidates are constrained. Only supergroups are tried, so the operations tested are ones that could be there rather than every operation the lattice permits, and the spurious acceptances at large tolerance never arise.
The output is a quantity, not a decision. A displacement in ångströms can be compared against the coordinate uncertainties the refinement already produced, which is the comparison that ought to be made and which a bare tolerance obscures.
The staircase remains the honest picture of what a naive detector does, and it is worth having for that reason. What the real tools show is that the way out is not a better threshold but a better question.
The same problem in the other direction
There is a mirror image of this difficulty which is worth naming, because it is the one that produces published errors rather than merely uncertainty.
Everything above concerns a structure that is symmetric and whose coordinates hide it. The commoner and more damaging case is a structure that is not symmetric and is described as though it were. Refining in a group that is too large forces atoms onto positions they do not occupy, and the fit is often acceptable — with fewer parameters and a residual only slightly worse, which looks like parsimony rather than error.
The symptoms are known and none of them is a bad residual. Unusually large thermal parameters, because an atom forced to a special position absorbs its real displacement into apparent vibration. Chemically implausible bond lengths, averaged between two real ones. And a difference map with structure in it near the atoms that were constrained.
The reason this matters here is that it inverts the direction of caution. A loose tolerance finds symmetry that is not there, and a strict one misses symmetry that is; but the consequences are not symmetric. Missing a symmetry gives a structure with too many parameters and a slightly overfitted description, which is recoverable. Imposing one that is absent gives coordinates that are wrong, and every distance computed from them is wrong with them.
Where the exactness stops
This whole essay is about a place where exactness stops, so the boundary needs stating with more than usual care.
The displacement here is known and uniform. Every atom was moved by up to the same amount, in a direction chosen by a seeded generator. Real errors differ per atom, are correlated between atoms bonded to one another, and are larger in some directions than others. A single tolerance is the wrong model of that even before anybody chooses a value for it.
The detector here is naive. It accepts an operation when every image is near some atom, which permits two atoms to map onto one. Structure-comparison software is more careful, and the count of accepted operations at large tolerance would be lower with a better matcher — the staircase’s first two steps would be unchanged, since they are where the matching is unambiguous.
And nothing here is a claim about a real structure. The pattern is generated, the displacement is synthetic, and the numbers are properties of this experiment. What generalises is the shape of the curve, not the values on its axes.
One further contrast is worth stating because it explains why this essay can be exact about an inexact question. The staircase is computed on a pattern this site generated, so the true answer is known throughout — the group is p4m and the displacement is a number that was chosen. That is what makes the plot a measurement of the detector rather than of the structure, and it is the only arrangement in which a threshold’s behaviour can be studied at all. A measured structure supplies no such ground truth, and the staircase it would produce could not be labelled.
Where the ladder goes next
The kind of extra symmetry that no care removes because it belongs to the measurement rather than the object is the symmetry diffraction adds.
The kind that a careless motif produces, and that the round trip refuses outright, is the motif must be a comma.
The reason the exact question is answerable at all — integer matrices in the lattice basis, and rational translations — is what a symmetry actually is, and the machinery that decides it is the orbit and its detector.
What the pictures here cannot show. The two panels of the second figure are the whole difficulty and they cannot demonstrate it: a reader who could see the difference between them would not need the staircase, and a reader who cannot has only the site’s word for it that the right-hand pattern has one symmetry operation. That claim is the output of an exact computation on coordinates the figure prints nowhere, and no drawing of a nearly-symmetric pattern can establish how nearly.
What this makes readable
Essays that name this one as a prerequisite.
What links here
The 8 essays that link to this one and share the most of its objects, of 50 that link here.
The objects this essay names
Each one links to every other essay that touches it.
Accidental symmetryPseudosymmetryRefinementRound tripSite symmetrySpace group determinationTolerance