Operations

One crystal, and sixteen coordinate lists

Two structure reports can disagree in every number and describe the same arrangement of atoms, because a space group does not fix its own origin or its own axes. How many genuinely different lists there are is the index of the group in its Euclidean normaliser — a number, computable, and the thing a structural database has to divide out before it can say two entries are the same compound.

Assumes The same pattern, described twice and Two origins for one group.

Two coordinate lists for one structure can disagree in every number and describe the same arrangement. That essay makes the point in the plane and reads the answer off pictures. This one asks the question in space, where the answer is a number that a structural database has to know.

The number matters because the database’s job is to answer have these two people determined the same structure?, and two determinations of one compound routinely arrive with entirely different coordinates.

What may be changed without changing anything

Fix a space group G. Which motions of space may be applied to a description of a structure in G without changing what is described?

A motion n may be applied when n G n⁻¹ = G — when conjugating every operation of the group gives the group back. Such motions form a group containing G, called its Euclidean normaliser N. Applying an element of G itself changes the coordinate list by permuting equivalent positions, which is no change at all; applying an element of N outside G gives a genuinely different list describing the same crystal.

So the number of descriptions is the index |N : G|, one for each coset.

P222: 16 descriptions of one structure. P222 has 4 operations in a cell. 8 origins leave every one of them exactly where it was, and 8 linear parts of the lattice's holohedry normalise the group, so its Euclidean normaliser has 64 elements per cell and the index is 16. That index is the number of coordinate lists that describe one and the same arrangement of atoms. Each was applied to a motif and the resulting point sets compared: the numbers differ and the sets are identical, which is the check that makes the count mean anything.
Fig. 1 The count for one group, in the four quantities it comes from. Four operations in the cell; eight origins that leave every one of them exactly where they were; eight linear parts of the lattice’s holohedry that normalise the group; sixty-four elements in the normaliser per cell; and sixty-four divided by four is sixteen descriptions. The index is a quotient, not a product: the group’s own four operations are among the sixty-four and change no description at all.

Computed rather than looked up

The normaliser is tabulated in the International Tables, and tabulated numbers are exactly what this collection declines to use. So it is searched for.

The linear parts can only come from the holohedry of the lattice — nothing else maps the lattice to itself, and a motion that does not is not a symmetry of anything the group describes. That is at most forty-eight candidates.

The translations are the origins that leave every operation exactly where it was, and they satisfy a linear condition: a translation t conjugates (A, a) to (A, a + (I − A)t), so t normalises exactly when (I − A)t is a translation of the group for every A.

That last phrase does real work. On a primitive lattice, a translation of the group means an integer vector. On a centred one it means an integer vector or a centring vector, and requiring integrality instead throws away the centring translations themselves — which made Fm3̅m come back with two origins where it has four, and an index of one half. An index of one half is not a number an index can be, and that impossibility is what caught it.

21 groups with a countable number of descriptions, 24 with a sliding origin. Space groups whose Euclidean normaliser has no free direction, with the number of origins that leave every operation where it was, the index of the group in its normaliser, and the number of genuinely different coordinate lists that describe one structure. P222 has 16 of them and Im-3m has 1: the most symmetric group is already its own normaliser and there is nothing to be done to a description of it. The other 24 groups are normalised by a continuum of translations — their origin slides — and are not in this table, because a count read off a grid would be a fact about the grid.
Fig. 2 The groups whose normaliser has no free direction, with their origin counts and their indices. The most symmetric group here is already its own normaliser — there is nothing to be done to a description of it — and the least symmetric have sixteen.

How the search is bounded

The linear parts come from a list of at most forty-eight; the translations come from a grid; and the grid’s fineness is the one parameter that could quietly decide the answer.

It cannot be chosen once and left. Fddd’s two conventional origins are an eighth of a cell apart, so a grid of sixths cannot represent the motion between them, and the search over a grid of sixths returns half the normalising motions and an index of one half. A grid of twelfths represents halves, thirds, quarters and sixths and still not eighths.

So the grid is raised until the index is a whole number at least one, and the grid that worked is reported. That is not a fudge: the index is an index, so a non-integer is a proof that the search was incomplete, and the machinery uses the impossibility as its own stopping rule. It is the same style of check as an odd coincidence index or a count that must come to seventeen — an arithmetic constraint used to detect a computational failure.

And the count of origins is checked separately, by refinement. Every group with no free direction has its origin count computed at twelfths and again at twenty-fourths, and the two are required to agree. A group whose count moved would be a group whose origins are not what the grid says.

Fddd: 2 descriptions of one structure. Fddd has 32 operations in a cell. 8 origins leave every one of them exactly where it was, and 8 linear parts of the lattice's holohedry normalise the group, so its Euclidean normaliser has 64 elements per cell and the index is 2. That index is the number of coordinate lists that describe one and the same arrangement of atoms. Each was applied to a motif and the resulting point sets compared: the numbers differ and the sets are identical, which is the check that makes the count mean anything.
Fig. 3 The group that forced the escalation. Its two conventional origins are an eighth of a cell apart, which is why the Tables print it twice and why a grid of sixths reports an impossible index for it.

The two normalisers, and which one a report needs

There is a second normaliser, larger than the one computed here, and the difference is worth stating because reports use both.

The Euclidean normaliser allows motions: rotations, reflections and translations. It is what has been computed, and its cosets are the descriptions of one structure in one setting with one choice of axes.

The affine normaliser allows any change of basis preserving the group — including relabelling the axes and rescaling them. For a triclinic group that is an enormous set, and the index is infinite; for a cubic group the two coincide.

Which one is wanted depends on the question. Comparing two determinations made in the same setting needs the Euclidean one. Comparing across settings — one report in Pnma and another in Pbnm, which are the same group on relabelled axes — needs the affine one, and is what the six ways of naming one group is about.

When the count is not a number

Some groups are normalised by a continuum of translations, and for those the question has no numerical answer.

P1 has no operations but the identity, so every translation normalises it: the origin may be put anywhere, in all three directions. P2 is normalised by any translation along its own axis and by half-cell shifts across it, so its origin slides in one direction. In both cases the count of origins depends entirely on how finely the search grids the cell, and a number read off a grid would be a fact about the grid.

Detecting that is exact linear algebra rather than a comparison of two grids. The condition (I − A)t = 0 over the reals defines the directions every operation fixes, and the dimension of that space is three minus the rank of the stacked matrices I − A — computed in integers, once, with no grid anywhere.

The two-grid version fails here and it is instructive how. Running the search at thirds and again at sixths, six of these groups appear to gain origins on refinement and would be reported as having a free direction. They do not: a grid of thirds cannot represent a half, so refining to sixths adds origins that were always there and were unrepresentable. A refinement test detects unrepresentability as readily as continuity, and the two look identical in the ratio.

Twenty-four of the forty-five groups built here slide. The other twenty-one have a count. Which is which is decided by whether any direction is fixed by the whole point group — the same quantity that decides whether a class is polar, arriving in a completely different question.

P2_12_12_1: 16 descriptions of one structure. P2_12_12_1 has 4 operations in a cell. 8 origins leave every one of them exactly where it was, and 8 linear parts of the lattice's holohedry normalise the group, so its Euclidean normaliser has 64 elements per cell and the index is 16. That index is the number of coordinate lists that describe one and the same arrangement of atoms. Each was applied to a motif and the resulting point sets compared: the numbers differ and the sets are identical, which is the check that makes the count mean anything.
Fig. 4 The most common space group in small-molecule crystallography, and the commonest for proteins: sixteen descriptions of one structure. Its four quantities are the same four as P222’s above — four operations, eight origins, eight linear parts, sixty-four — and that is not a coincidence to be explained away. The two groups have the same order on the same lattice, so they sit the same distance below the same ceiling, and the index cannot tell a group of screws from a group of rotations. What differs is everything else about them.

Using the answer

A count of descriptions that has never been shown to describe the same thing is a count of nothing, so the answer is used rather than reported.

A motif is placed in the cell and its orbit under the group generated — a set of points, which is the structure. Then each coset representative is applied to the motif and the orbit regenerated. The coordinate lists differ; the point sets are required to be identical, as sets, and they are for every group and every coset here.

That is the check that makes the number mean what it says. Without it the count is the size of a group, and with it the count is the number of ways of writing down a crystal.

P2₁2₁2₁, in the two diagrams the Tables print. Space group P2₁2₁2₁, number 19, projected down c on a primitive orthorhombic cell. The symmetry elements drawn: 8 2₁ screw axes. 4 general positions, the orbit of one point, each labelled with its height along c and marked with a comma where the operation that produced it reversed handedness.
Fig. 5 The structure the descriptions are of: one orbit of a general point under the group. Every one of the sixteen coordinate lists generates exactly this set of points from a different starting motif, and no measurement distinguishes them.

What a database does with it

The practical consequence is the reason the International Tables print these numbers at all.

Two structure reports of one compound are compared by reducing both to a standard description. The reduction applies each coset representative in turn, computes a canonical form of the resulting coordinate list — sorted, with a rule for choosing among equal candidates — and keeps the smallest. Two structures agree when their canonical forms do.

Skipping that step produces false negatives, and they are common. A search that compares raw coordinates finds two determinations of one compound to be different structures in fifteen cases out of sixteen for a P2₁2₁2₁ crystal, which is the failure this arithmetic exists to prevent. The same failure in reverse — two different structures called the same — is what a wrong equivalence always produces.

And skipping it in the other direction produces a subtler error. A structure and its enantiomorph have coordinate lists related by an inversion, which is not in the normaliser of a Sohncke group. So the two are genuinely different structures, and a comparison that quotiented by too much would call a left-handed crystal and a right-handed one the same. The normaliser is exactly the right amount to divide out: no more and no less.

19 of the 45 built here. The space groups this site builds whose operations all preserve orientation. There are 19 of them among the 45 it constructs, and the International Tables record 65 among the two hundred and thirty. That number is quoted rather than enumerated, as two hundred and thirty is: what is computed here is the criterion, on every group available. Four of these have an enantiomorphic partner — a distinct group that is their mirror image — and the rest are their own.
Fig. 6 The groups a structure of a single hand may sit in. For these the inversion is outside both the group and its normaliser, so the two enantiomorphs stay distinct through every reduction — which is what a database needs, since they are different compounds.

The origins, and where they were met before

The translations in the normaliser are the equivalent origins, and this collection has met them from another direction.

Two origins for one group is about the twenty-four space groups the Tables print twice, once with the origin at a centre of symmetry and once at a point of higher site symmetry. That is a choice between two conventions, and it is a subset of the freedom counted here, in the same way a setting is a subset of the relabellings available: the normaliser’s translations include those two and usually more.

The difference between the two accounts is what is being counted. The Tables count settings a crystallographer might reasonably use. The normaliser counts every origin that leaves the operations where they were, including ones no convention would choose. The first is a decision about presentation and the second is a fact about the group.

P-1: 8 descriptions of one structure. P-1 has 2 operations in a cell. 8 origins leave every one of them exactly where it was, and 2 linear parts of the lattice's holohedry normalise the group, so its Euclidean normaliser has 16 elements per cell and the index is 8. That index is the number of coordinate lists that describe one and the same arrangement of atoms. Each was applied to a motif and the resulting point sets compared: the numbers differ and the sets are identical, which is the check that makes the count mean anything.
Fig. 7 The simplest non-trivial case: P1̅, with eight equivalent origins — the eight inversion centres of the cell — and an index of eight. Anybody who has moved a structure from one centre to another has used this number without computing it.

What the number is a property of

One clarification, because the phrase number of descriptions invites a misreading.

The index is a property of the group, not of the structure. A structure with high site symmetry may be carried onto itself by some of the normaliser’s cosets, in which case fewer than |N : G| distinct coordinate lists appear — the extra ones coincide. So the index is an upper bound on how many lists a particular structure has, exact when the structure occupies general positions and smaller otherwise.

A database has to handle that, and does, by applying all the cosets and deduplicating rather than assuming the count. The figures here place a motif in a general position for exactly that reason.

Pm-3m: 7 distinct site symmetries, up to order 48. Every distinct site symmetry of Pm-3m, found by taking every point of a grid of twelfths and asking which of the group's 48 operations leave it exactly where it is. The second column is the multiset of operation types — a rotation of each order, a mirror, an inversion — which names the site's point group without any table being consulted. The third is how many grid points have that symmetry, which is why the general position dominates: almost every point is fixed by nothing.
Fig. 8 Where the exception lives: a group’s distinct site symmetries. A structure whose atoms sit at the highest of these is invariant under more of the normaliser than a structure in general positions, and it has fewer descriptions.

Why the index is what it is

Reading the table above, the pattern is that the most symmetric groups have the smallest index, which is the opposite of what a first guess would suggest.

The reason is that the normaliser is bounded above by the holohedry’s motions and below by the group itself. A group that already contains most of what the lattice permits has little room above it, so the index is small; a group with few operations sits inside the same ceiling with a great deal of room, so the index is large.

Im3̅m has an index of one. It is the full symmetry of its lattice, its normaliser is itself, and there is exactly one way to write down a structure in it once the cell is fixed. P222 has sixteen, because a group of order four sits inside a normaliser of order sixty-four.

That is a useful thing to know before reading a structure report. A triclinic structure has an origin that can be anywhere, and two determinations of one will agree in nothing at all until both are reduced; a cubic structure in a high-symmetry group has almost no freedom, and two determinations should agree immediately.

Read the table at the head of this essay that way and the correlation is negative, which is not an accident: it is the statement that the normaliser has a ceiling and the group is climbing towards it. The ceiling is the same for every group on one lattice — the holohedry’s motions, and nothing else can normalise anything — so the index measures the gap between a group and the most symmetric group its own lattice allows. Two groups with the same index are two groups the same distance below their ceilings, which is why the number tracks the group’s order rather than its symbol.

P2_12_12_1: 1 distinct site symmetries, up to order 1. Every distinct site symmetry of P2_12_12_1, found by taking every point of a grid of twelfths and asking which of the group's 4 operations leave it exactly where it is. The second column is the multiset of operation types — a rotation of each order, a mirror, an inversion — which names the site's point group without any table being consulted. The third is how many grid points have that symmetry, which is why the general position dominates: almost every point is fixed by nothing.
Fig. 9 And the extreme case: a group with a single kind of site. Every point of the cell has trivial site symmetry, so no structure in this group is ever invariant under any coset, and every one of its sixteen descriptions is genuinely distinct.

What is owned, and what is not

Owned: the free-direction test by rank, the normalising translations under the group’s own translation lattice, the linear parts by search over the holohedry with the grid raised until the index is a whole number, the index, and the check that every coset gives the same point set.

Not owned: the affine normaliser, which allows changes of basis as well as motions and gives a different and larger count; the tabulated normalisers of the Tables, which are not consulted; and any claim about a particular database’s algorithm.

The refusals: a group with a sliding origin is refused a finite count, and a motion outside the normaliser is refused — the second checked by conjugating with a translation of a fifth of a cell and requiring the operation set to come back different.

Pnma: 8 descriptions of one structure. Pnma has 8 operations in a cell. 8 origins leave every one of them exactly where it was, and 8 linear parts of the lattice's holohedry normalise the group, so its Euclidean normaliser has 64 elements per cell and the index is 8. That index is the number of coordinate lists that describe one and the same arrangement of atoms. Each was applied to a motif and the resulting point sets compared: the numbers differ and the sets are identical, which is the check that makes the count mean anything.
Fig. 10 A third group, chosen because it is the second commonest in inorganic structure reports. Eight descriptions, which is why the same compound appears in the literature with coordinates that look unrelated and reduce to the same thing.

At the other end the freedom has almost run out. Pm3̅m has two descriptions and Im3̅m has one, so a cubic structure in a high-symmetry group determined twice will agree on nearly everything with no reduction applied at all — and a crystallographer who has only ever worked in such groups may reasonably never have met this arithmetic. The people who meet it are the ones working in P2₁2₁2₁ and P2₁/c, which is most of small-molecule crystallography and all of protein crystallography, and where the number is sixteen and eight.

The same ambiguity, during the determination rather than after it

The count above is presented as a problem for a database comparing finished structures. It is a problem during the determination too, and there it has a name and a standard treatment.

A phase determination fixes an origin by assigning phases to a few reflections, and the origins available to be fixed are exactly the translations in the normaliser — so the number of ways to fix the origin is the count on this page. In protein crystallography the group of allowed shifts is called the Cheshire group, and knowing it is what makes several standard operations well posed.

Molecular replacement is the clearest case. A search for where a known fragment sits in an unknown cell returns a rotation and a translation, and it returns them modulo the normaliser: several solutions with different translation vectors are one solution described several ways, and a search reporting them as distinct is reporting the same answer repeatedly. Programs reduce every candidate to a canonical coset before ranking, for exactly the reason a database does.

Averaging two maps requires the same reduction. Two independent phase determinations of one crystal produce two maps, and comparing or combining them means putting them on a common origin first. Without that they disagree everywhere, by an amount that carries no information about either.

So the index is not a bookkeeping number appended after the work. It is a parameter of the work, and a determination that has not computed it is one in which several later steps are ill-defined.

The half that a chiral structure loses

There is a case where the count is halved, and it is worth separating because it is the one an inexperienced comparison gets wrong in the more damaging direction.

For a group containing no operation that reverses handedness, the inversion is not in the Euclidean normaliser — conjugating by it produces the mirror-image group, which for an enantiomorphic pair is a different group and for the rest is the same group described in a left-handed frame. Either way the inversion is not a motion that may be applied while claiming nothing has changed.

That has a consequence for the comparison. Two coordinate lists related by an inversion describe a structure and its mirror image, which for a chiral compound are two different substances. A database routine that includes the inversion among its coset representatives — because it was written for the centrosymmetric case, where the inversion is in the group anyway and costs nothing — will report the two enantiomers of a compound as one entry.

That is a false positive of the worst kind, because it merges two records rather than duplicating one, and nothing downstream can recover the distinction. The safeguard is the same one the essay’s search uses everywhere: build the normaliser from the group’s own operations rather than from a general list of motions, and let the group decide which of them belong.

What the sixteen look like as a list

It helps to see what “sixteen descriptions” means for one crystal, because the phrase suggests sixteen sets of numbers with something visibly in common and that is not what they are.

Take a single atom at (0.11, 0.23, 0.37) in P2₁2₁2₁. The eight origins available are the eight half-cell shifts — every combination of 0 and ½ in the three directions — so moving to each of them sends the atom to (0.11, 0.23, 0.37), (0.11, 0.23, 0.87), (0.11, 0.73, 0.37) and five more, none of which shares a digit with the others in every place. The eight linear parts then act on each of those, reversing pairs of axes; that is sixty-four triples, of which the group’s own four operations identify each with three others, leaving sixteen genuinely different coordinate lists.

Nothing about a list says which of the sixteen it is. There is no distinguished one, because the group provides no way of preferring an origin: all four are equally the origin, in the exact sense that the operation set is identical at each. So a canonical form has to be imposed — sort the atoms, apply every coset, and keep whichever list is smallest under some ordering — and the ordering is a convention with no mathematical content whatever. It has to be the same convention at both ends of a comparison and it does not have to be a good one.

And the number is not the number of coordinate lists a reader will meet. It is the number a general structure has. A structure with an atom on a special position is invariant under some of the cosets, and its list count divides sixteen rather than equalling it; a structure with several atoms in general positions has all sixteen. The count is therefore an upper bound with a divisor structure underneath it, which is what makes deduplication — apply everything, then compare — the right implementation rather than an inefficient one.

Where the ladder goes next

Sideways, to the other freedom a description has. One group, three symbols is about the axes rather than the origin: relabelling a, b and c gives another description of the same crystal, and the group of relabellings is what the affine normaliser adds to the Euclidean one.

Down, to what a description cannot change however it is written. A site’s symmetry is the same in every coordinate list, because it is the stabiliser of a point and conjugation moves the point and the stabiliser together — which is the sense in which some things about a structure are description-independent and worth reporting.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

CosetEquivalenceEquivalent originHolohedryIndexNormaliserOrigin choiceSetting