How it is known

A twin hides in the statistics

A twinned crystal scatters as two orientations at once and the detector cannot separate them. What arrives is a sum of two intensities — and adding two independent quantities narrows a distribution, which is a signature no model of the structure is needed to read.

Assumes Whether there is a centre is a statistic, A twin is a symmetry the lattice has and the crystal does not and Twenty-five of the thirty-two can twin, and seven cannot.

A twin is a symmetry the lattice has and the crystal does not, and this collection has treated it as a group-theoretic object: the twin laws available to a structure are the cosets of its point group in the lattice’s holohedry, the count is exact, and twenty-five of the thirty-two classes can twin while seven cannot.

That is the arithmetic of which twins are possible. This is the measurement of whether one is present, and it is a statistical argument rather than an exact one — which is why it sits here beside the test for a centre of symmetry rather than with the cosets.

The twin fraction, recovered from a moment and nothing else (12 atoms). A structure of 12 atoms twinned at each of 6 fractions, with the second moment of its intensity distribution measured and the fraction solved back out of it. The recovery is within a few hundredths as far as thirty per cent — 4 rows here — and 4 of the 6 fractions get a number at all. Beyond thirty per cent the relation flattens: the derivative of 2α(1−α) vanishes at a half, the two roots meet, and a small error in the moment becomes a large one in the fraction. Where the sampled moment falls below 1.5 the quadratic has no real root and the estimate refuses rather than clamping, which is why a nearly perfect twin is the hard case in practice rather than the easy one.
Fig. 1 A structure twinned at six different fractions, with the second moment of its intensity distribution measured and the fraction solved back out of it. The recovery is good to a few hundredths as far as thirty per cent and then refuses, because near a half the two roots of the relation meet and a sample of a few hundred reflections cannot tell them apart.

What a twinned crystal sends to the detector

A twinned crystal is two orientations of one structure sharing a lattice. If the twin law is a merohedral one — an operation of the lattice’s holohedry that is not in the crystal’s point group — then the two orientations’ reciprocal lattices coincide exactly, and the two reflections that land at the same place are h and h′ = hT.

The detector records their sum, weighted by how much of the crystal is in each orientation:

Iobs(h)=(1α)I(h)+αI(h)I_{\text{obs}}(h) = (1-\alpha)\,I(h) + \alpha\,I(h')

with α the twin fraction. Nothing about the measurement separates them: the spots are in the same places, and the intensities have already been added by the time anything is recorded. A merohedral twin moves no spot at all, which is what makes it dangerous.

Why the distribution narrows

Adding two independent quantities makes the result less variable, and that is the whole of the signature.

For an acentric structure the normalised intensities follow an exponential distribution, whose second moment ⟨I²⟩/⟨I⟩² is exactly 2. For a perfect twin the observed intensity is the average of two independent draws from that distribution, and the second moment falls to 1.5. In between,

I2/I2=22α(1α)\langle I^2\rangle / \langle I\rangle^2 = 2 - 2\alpha(1-\alpha)

which is a quadratic in α that can be solved backwards.

The derivation is three lines of algebra with no crystallography in it: expand the square of a weighted sum of two independent exponentials, use ⟨I⟩ = 1 and ⟨I²⟩ = 2, and collect. What makes it useful is that the answer needs no model of the structure — not the atoms, not the cell contents, not a single phase.

Twinning narrows the distribution: ⟨I²⟩/⟨I⟩² falls from 1.97 to 1.55. The cumulative distribution of normalised intensities for one structure, measured twice: as itself, and with a twin fraction of 0.3 mixed in. Twinning adds two independent intensities together, and adding two independent quantities makes the result less variable — so the twinned curve has fewer very weak and fewer very strong reflections, and rises more steeply through the middle. Nothing about the structure is used in reading this: the whole of the twin test is that the second moment of the distribution falls from two towards one and a half.
Fig. 2 The same structure measured twice: as itself, and with a twin fraction of 0.3 mixed in. The twinned curve has fewer very weak and fewer very strong reflections and rises more steeply through the middle — the distribution has narrowed, and the second moment has fallen from 1.97 to 1.55.

Two roots, and why the ambiguity is real

The relation is a quadratic and has two roots, α and 1 − α, and no measurement separates them.

That is not a defect of the estimator. A crystal twinned at thirty per cent and one twinned at seventy per cent scatter identically, because the two orientations are not labelled: calling one of them “the crystal” and the other “the twin” is a choice, and swapping the names exchanges the roots. The physical situations are mirror descriptions of one arrangement.

So the estimate reports both, and the convention is to quote the smaller. Anything else would be attaching a meaning to a label the experiment never assigned.

Where the estimate refuses

Near a half the relation flattens: the derivative of 2α(1 − α) vanishes at α = ½, so a small error in the moment becomes a large error in the fraction, and the two roots have converged.

Worse, the sampling scatter of a moment computed from a few hundred reflections is enough to push the measured value below 1.5, at which point the quadratic has no real root at all. The machinery here reports refused in that case rather than clamping to a half or returning a complex number, and the refusal is the honest output: this sample cannot resolve a fraction this close to a half.

That is a general property of the method rather than an artefact of the implementation, and it has a consequence worth knowing: a nearly perfect twin is the hard case, not the easy one. A crystal twinned at forty-eight per cent is almost indistinguishable from one twinned at fifty, and both are almost indistinguishable from a crystal in the higher-symmetry space group that the twin law would produce.

Twinning narrows the distribution: ⟨I²⟩/⟨I⟩² falls from 1.97 to 1.47. The cumulative distribution of normalised intensities for one structure, measured twice: as itself, and with a twin fraction of 0.5 mixed in. Twinning adds two independent intensities together, and adding two independent quantities makes the result less variable — so the twinned curve has fewer very weak and fewer very strong reflections, and rises more steeply through the middle. Nothing about the structure is used in reading this: the whole of the twin test is that the second moment of the distribution falls from two towards one and a half.
Fig. 3 A perfect twin, whose distribution is as narrow as twinning can make it. Its second moment is 1.5 and everything about it looks like a well-behaved crystal in a higher-symmetry group — which is exactly the confusion this test exists to raise and cannot always settle.

Where the relation comes from, in three lines

The relation is short enough to derive rather than quote, and deriving it makes both of its failure modes visible.

An acentric structure’s normalised intensities are exponentially distributed: ⟨I⟩ = 1 and ⟨I²⟩ = 2. The observed intensity is (1−α)I₁ + αI₂ with I₁ and I₂ independent draws from that distribution.

Expanding the square and taking expectations:

Iobs2=(1α)22+α22+2α(1α)1=22α(1α)\langle I_{\text{obs}}^2\rangle = (1-\alpha)^2\cdot 2 + \alpha^2 \cdot 2 + 2\alpha(1-\alpha)\cdot 1 = 2 - 2\alpha(1-\alpha)

with the cross term contributing ⟨I₁⟩⟨I₂⟩ = 1 rather than 2 because the two are independent. And ⟨I_obs⟩ = 1, so the ratio is the same expression.

The two failure modes are now visible. If the distribution is not exponential — a centric structure, whose second moment is 3 — the first line is wrong and everything after it is wrong. And if I₁ and I₂ are not independent — a twin law relating a reflection to itself, or to one whose intensity is correlated with it — the cross term is not 1, and the relation over-estimates the fraction.

The second failure has a name: a twin law that is already an operation of the crystal’s Laue group relates each reflection to one it equals exactly, so twinning by it changes nothing at all and is undetectable. Which laws are like that is decided by the coset arithmetic rather than by any statistic.

The trap: a centre of symmetry looks like the opposite

The relation above is for an acentric distribution, and the centre test established what a centrosymmetric one does instead: its second moment is 3 rather than 2, because a centric structure’s intensities have a much longer tail.

So a centrosymmetric structure, fed to the twin estimator, produces a moment of about 3 — above the untwinned acentric value, in the direction opposite to twinning — and a formula applied without thinking returns a negative twin fraction.

The machinery refuses that outright: a moment above 2 is reported as “this looks centric, not twinned”, with no number attached. Refusing is the point. A negative fraction would be recognised as nonsense by any reader; a fraction of zero-point-something, produced by clamping, would not.

That failure mode is not hypothetical. It is the commonest misuse of intensity statistics in practice, and it happens because the two tests use the same measured quantity to answer different questions: the moment says something is unusual about the distribution, and deciding whether the something is a centre or a twin needs more than the moment.

The twin fraction, recovered from a moment and nothing else (20 atoms). A structure of 20 atoms twinned at each of 6 fractions, with the second moment of its intensity distribution measured and the fraction solved back out of it. The recovery is within a few hundredths as far as thirty per cent — 3 rows here — and 5 of the 6 fractions get a number at all. Beyond thirty per cent the relation flattens: the derivative of 2α(1−α) vanishes at a half, the two roots meet, and a small error in the moment becomes a large one in the fraction. Where the sampled moment falls below 1.5 the quadratic has no real root and the estimate refuses rather than clamping, which is why a nearly perfect twin is the hard case in practice rather than the easy one.
Fig. 4 The ladder again with a larger structure, where the scatter is smaller and the recovery holds a little further. Nothing about the relation changes; what changes is the precision of the sample mean, which is the only thing standing between the measurement and the answer.

What the test is for, in a real determination

Three uses, and they are the reason this is a standard part of processing a data set rather than a curiosity.

Before solving. A twinned data set refined as though it were untwinned gives a poor fit and an oddly high residual, and the structure may not solve at all. Detecting the twin first turns an intractable problem into an ordinary one with an extra parameter.

As a diagnosis of the wrong space group. A crystal assigned to a space group larger than its own produces intensity statistics that look twinned, because merging inequivalent reflections is arithmetically the same operation as a twin averaging them. The two are genuinely confusable and the statistics say so.

As a warning about pseudo-symmetry. A crystal whose lattice is nearly of a higher type can twin by pseudo-merohedry, which puts reflections nearly but not exactly on top of one another. The statistics are then intermediate and messy, and the moment is the first sign that something is going on.

The other statistic, and why two are better than one

The moment is one number and a data set contains a distribution, so a second test using more of it is available and is standard: the cumulative distribution, compared against the two closed forms the centre test uses.

An untwinned acentric data set follows 1 − e⁻ᶻ; a centric one follows an error function; and a twinned acentric one follows neither, sitting inside both with a shape of its own. Plotting all three together separates the three situations by eye, where the moment alone gives one number that two of them could produce.

That is the practical reason a processing pipeline computes several statistics rather than one: the moment says the distribution is unusual, the cumulative says in which direction, and the two together distinguish a twin from a centre from a merging error.

This collection computes the moment because it is the one whose algebra is closed-form and whose failures are legible. A reader taking the method to real data should compute both.

The three moments the centre test uses make the same point in numbers rather than in a picture. An acentric distribution has a second moment of two and a centric one of three, and a twinned acentric structure sits between the two — below two, because twinning narrows, and never above it. So a measured moment of 1.6 is a twin, a measured moment of 2.9 is a centre, and the interval between two and three is where the two tests part company rather than overlap. That is the whole reason the sign of the departure carries as much information as its size.

What is measured, and how honestly

Everything in this essay is a measurement, and it is worth listing what is being averaged over.

The structures are generated from a stated seed with a stated number of atoms; the intensities are computed by the same structure-factor sum every diffraction figure in this collection uses; the moment is a sample mean over a block of reflections. Change the seed and the recovered fractions move by a few hundredths.

That is not a weakness to be hidden — it is what the method is. A twin fraction estimated from intensity statistics has an uncertainty of a few per cent, and a determination that quotes it to four decimal places is quoting the arithmetic rather than the measurement. The figures here report the recovered value beside the true one so the size of the disagreement is visible rather than described.

N(z): the fraction of reflections weaker than z. The cumulative distribution of normalised intensities, measured on two structures built from the same atoms — one with an inversion centre, one without — and drawn against the two closed forms, 1 − e^(−z) without a centre and erf(√(z/2)) with one. The curves are furthest apart at small z, which is the useful end: a centrosymmetric structure has far more nearly-absent reflections, because its structure factor is a single real number that can pass through zero rather than a complex one that rarely does.
Fig. 5 The centre test, whose machinery this one borrows: the cumulative distribution of normalised intensities against the two closed forms. A twinned acentric distribution sits between them — narrower than acentric, and for a quite different reason from centric — which is why the two tests have to be read together rather than one at a time.
p3, twinned. p3 twinned by a rotation. To the left of the composition line the motif sits where p3 puts it; to the right every copy has been carried over by the twin law, which is one of the 3 operations the hexagonal lattice has and p3 does not. 24 images on the left, 24 on the right, and the lattice runs through the line unbroken — which is exactly why a twinned crystal looks like a single one.
Fig. 6 The object being measured, drawn: two orientations of one structure sharing a lattice, related by an operation the lattice has and the crystal does not. To the left of the composition line the motif sits where p3 puts it and to the right every copy has been carried over by the law. Everything in this essay is about detecting that arrangement from intensities alone, without ever seeing the two orientations separately — and the figure refuses to be drawn unless the law really does move the point set, since two halves that agreed would be a picture of no twin at all.

Why the moment and not the whole distribution

The test above uses one number — the second moment — and the distribution contains a great deal more. Using more of it is possible and is what the standard tests do, and it is worth saying why the moment is the right thing to lead with.

The moment is a sample mean, so its uncertainty falls as the square root of the number of reflections and is easy to state. A test based on the shape of a cumulative distribution is more sensitive and its uncertainty is a harder thing to write down, which matters when the answer is going to be quoted with an error bar.

The moment also has the cleanest relation to the fraction: a quadratic, invertible in closed form, with the two roots and the flat region near a half all visible in the algebra rather than discovered by experiment. The behaviour of a shape test near a perfect twin has to be measured.

And the moment is what fails safely. Fed a centrosymmetric structure it returns a number above 2, which is outside the range the relation covers and is refused. A shape test fed the same data returns a curve that does not match either closed form, which requires a judgement rather than triggering a refusal.

None of that says the moment is the best test. It says it is the one whose failures are legible, which is the property this collection optimises for.

Twinning narrows the distribution: ⟨I²⟩/⟨I⟩² falls from 1.97 to 1.72. The cumulative distribution of normalised intensities for one structure, measured twice: as itself, and with a twin fraction of 0.15 mixed in. Twinning adds two independent intensities together, and adding two independent quantities makes the result less variable — so the twinned curve has fewer very weak and fewer very strong reflections, and rises more steeply through the middle. Nothing about the structure is used in reading this: the whole of the twin test is that the second moment of the distribution falls from two towards one and a half.
Fig. 7 A lightly twinned crystal, at fifteen per cent. The two curves are close and the difference between them is a fraction of a per cent in the moment — which is why a light twin is detectable only with a good sample, and why the recovered fraction at the low end of the ladder is the least reliable part of it.

The two halves of twinning, side by side

It is worth putting the exact and the statistical accounts next to each other, because this collection now has both and they answer different questions.

Which twin laws are possible is exact. The available operations are the cosets of the crystal’s point group in the holohedry of its lattice, the count is an index, and quartz has exactly three twin laws because its lattice permits exactly three cosets. No measurement is involved and no tolerance appears.

Whether a crystal is twinned, and by how much is statistical. It is read from the distribution of measured intensities, it has an uncertainty, and it refuses near a half.

Both are needed. The first says what to look for; the second says whether it is there. A determination that had only the first would have no way to know, and one that had only the second would have no candidate laws to test.

The exact half has no uncertainty in it at all, and it is worth being precise about what that means beside a statistic. The twin laws available to a crystal are the cosets of its point group in the holohedry of its lattice; the number of them is an index, which is a quotient of two integers; and running the partition over all thirty-two classes produces the same answer every time it is run, on every crystal of every substance in that class. Nothing is averaged and no sample size appears. The statistical half then decides which of those laws, if any, a particular specimen has taken up, and it answers with a number and an error bar. A determination needs both, and confusing the two is how a twin fraction comes to be quoted to four decimal places.

p3, single and twinned. Left, the diffraction pattern of a single crystal of p3. Right, the same crystal twinned, with 50 per cent of it in one orientation. Not one spot has moved — the twin law is a symmetry of the lattice, so the two reciprocal lattices lie exactly on top of one another — and 72 of the 81 reflections drawn have changed intensity. At a fifty-fifty twin the pattern acquires the full symmetry of the lattice's point group and is indistinguishable from a crystal that genuinely has it.
Fig. 8 Why the detector cannot separate them: the two orientations’ reciprocal lattices coincide exactly for a merohedral twin, so the spots land on top of one another and what is recorded is already a sum. Everything statistical in this essay exists because this picture makes the direct approach impossible.

Undoing the sum, and why it gets impossible at the same place

The estimator returns a fraction. What a determination wants next is the untwinned intensities, and recovering them is a small piece of linear algebra whose failure is the same failure, arriving for a visibly different reason.

The twin law pairs reflections. For a pair h and h', the two observations are

I_obs(h)  = (1−α) I(h)  + α I(h')
I_obs(h') = (1−α) I(h') + α I(h)

which is two equations in two unknowns, and solving them gives the true intensities in terms of the measured ones. The matrix to invert has determinant (1−α)² − α² = 1 − 2α.

That determinant vanishes at a fraction of a half. As the twinning approaches perfect, the two equations approach the same equation, and the inversion amplifies the measurement errors by a factor going as 1/(1 − 2α) — twofold at a third, tenfold at 0.45, and unbounded at a half. A perfectly twinned data set contains no information about which of the two orientations any intensity belonged to, and detwinning it produces numbers with the right mean and no meaning.

So the estimator’s flat region and the detwinning’s singular matrix are one fact seen twice. It is worth stating that way because the two look like separate practical difficulties and are not: near a half the observation I_obs(h) and I_obs(h') are the same quantity, so nothing computed from them can distinguish the two orientations, whatever the method.

The modern practice does not detwin at all. A refinement takes the twin fraction as a parameter and compares its model against the summed intensities, which is the quantity actually measured — so the badly conditioned inversion never happens, and the errors stay where the experiment put them. Detwinning survives as a way of making a map to look at, with the caution that the map’s quality falls off exactly as the fraction approaches a half.

The statistic that does not care about the rest of the data

The moment test has a failure mode the essay names above: anything else that makes intensities more uniform imitates twinning. There is a standard second test designed around that failure, and its design is worth following because it is a general trick.

The trouble is that the moment is a global statistic. It compares each intensity against the mean of its resolution shell, so anything distorting that mean — anisotropic diffraction, a pseudo-translation putting half the reflections systematically weak, a bad scaling — distorts the moment.

The repair is to compare intensities only against their neighbours. Take pairs of reflections close together in reciprocal space and not related by any symmetry, and form the quantity |I₁ − I₂| / (I₁ + I₂) for each pair. Its distribution is known for untwinned data and for twinned, and because both members of a pair are at nearly the same resolution and nearly the same direction, everything the global statistic was vulnerable to cancels between them.

That is the same move as several elsewhere in this collection: make the comparison local, so the systematic error is common to both sides of it. A twinned data set has fewer near-equal-and-opposite pairs than an untwinned one, for the same reason its moment is smaller — averaging two independent numbers pulls both towards the middle — and the test reads that directly, without ever needing a shell mean.

Where this goes

The natural next rung is the other twin test — the one that uses pairs of related reflections rather than the distribution as a whole, and which localises the twin law as well as the fraction. That test needs the candidate law in hand, which is where the exact half of the subject supplies it, and it is more than one essay’s worth of work.

The nearer neighbour is the redundancy of a data set, whose merging residual is the other early sign that a crystal is not in the class it was assigned — and whose orbit arithmetic is what a twin quietly interferes with.

Both tests answer the same question and neither is a substitute for the other. The moment is one number with a closed-form relation to the fraction and a legible failure; the local test is a distribution with no such relation and a much narrower set of things that can fool it. A processing pipeline runs both, and the case worth attention is the one where they disagree.

What this makes readable

Essays that name this one as a prerequisite.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Acentric distributionIntensity statisticsMerohedral twinNormalised intensitySecond momentTwin fractionTwinning