Planet Musings

July 23, 2026

Terence TaoA digestion of the Jacobian conjecture counterexample

The notorious Jacobian conjecture can be formulated concretely over the complex numbers as follows.

Conjecture 1 (Jacobian Conjecture) Let {F:{\bf C}^n \rightarrow {\bf C}^n} be a polynomial map in {n} complex variables, whose Jacobian {\mathrm{det} DF} is a non-zero constant. Then {F} is invertible (with polynomial inverse).

The condition that the Jacobian {\mathrm{det} DF} is non-zero is equivalent to {F} being locally invertible. (The implication of local invertibility from non-vanishing Jacobian follows from the inverse function theorem; the converse implication can be derived from the Weierstrass preparation theorem, but is omitted here.) Also, from the fundamental theorem of algebra, once the Jacobian polynomial {\mathrm{det} DF} is non-zero, it must be constant. So the hypothesis “Jacobian {\mathrm{det} DF} is a non-zero constant” can be replaced with “{F} is locally invertible”. So the Jacobian conjecture can be viewed as an assertion that local invertibility implies global invertibility. The complex numbers can be easily replaced with other fields of characteristic zero by the Lefschetz principle, but I prefer to work in the concrete setting of the complex numbers.

It was recently shown (using the Fable AI) that the conjecture is false in three dimensions (and thus in higher dimensions as well):

Theorem 2 (Counterexample to conjecture) There exists a polynomial {F : {\bf C}^3 \rightarrow {\bf C}^3} which has non-zero constant Jacobian, but is not invertible.

The conjecture remains open in two dimensions, and is easy to establish in one dimension.

The example can be stated completely explicitly: one can take

\displaystyle  F(z_1,z_2,z_3) = \Big((1+z_1 z_2)^3 z_3 + z_2^2 (1+z_1z_2) (4+3z_1z_2), \ \ \ \ \ (1)

\displaystyle  z_2 + 3 z_1 (1+z_1z_2)^2 z_3 + 3 z_1 z_2^2 (4+3z_1z_2),

\displaystyle 2 z_1 - 3 z_1^2 z_2 - z_1^3 z_3\Big)

and one can verify by a brief calculation that

\displaystyle  \mathrm{det} DF = -2

and

\displaystyle  F(0,0,-1/4) = F(1,-3/2, 13/2) = F(-1,3/2,13/2)

\displaystyle  = (-1/4,0,0).

While this is an extremely quick verification, the construction presented in this fashion appears like a massive miracle. The polynomial {F} has degree seven, so a priori the Jacobian {\mathrm{det} DF} ought to be a polynomial in three variables of degree as large as {3 \times 6 = 18}, so the fact that all non-constant coefficients of this polynomial vanish looks like a massive cancellation involving {\binom{18+3}{3}-1 = 1329} equations, which is much larger than the {3 \times \binom{7+3}{3} = 360} degrees of freedom for a generic degree seven polynomial map of three variables. So finding such a polynomial looks highly unlikely to be located by brute force.

The example has since been retroactively explained in more geometric terms. As a “digestion” exercise to myself, I sought to write this explanation with relatively little use of algebraic geometry, in a manner that minimizes the amount of “miracles” required, although there are still a few places where some remarkable phenomena occur.

It is convenient to use the local injectivity formulation, and to generalize the domain {{\bf C}^3} to an equivalent affine variety. Namely, we will show

Theorem 3 (Counterexample, reformulated) There exists an affine variety {X \subset {\bf C}^5} that is isomorphic to {{\bf C}^3} by polynomial changes of variable, and a polynomial map {F : X \rightarrow {\bf C}^3} which is locally injective, but not globally injective.

Clearly one can get from Theorem 3 to Theorem 2 by composing with the isomorphism {X \cong {\bf C}^3} and using the previously mentioned fact that local injectivity implies non-zero constant Jacobian. Our objective is now to find data {X}, {F : X \rightarrow {\bf C}^3} that obeys three separate properties:

  • (a) {F} is locally injective on {X}.
  • (b) {F} is not globally injective on {X}.
  • (c) {X} is isomorphic to {{\bf C}^3} by polynomial changes of variable.
The advantage of splitting the problem in to these three components is that we can build towards each of them separately.

(A pedantic remark: strictly speaking, in the arguments below, we not only replace the domain {{\bf C}^3} of {F} by an equivalent variety {X}, but also replace the range {{\bf C}^3} of {F} by an equivalent variety {V}. But the equivalence between {V} and {{\bf C}^3} is a boring linear isomorphism ({V} will just be a hyperplane in a four-dimensional vector space {\mathrm{Sym}^3({\bf C}^2)}), so we do not highlight this aspect of the construction.)

It turns out that {F} and {X} can be built out of the operation of multiplication of low degree polynomials. Namely, consider the following three simple affine spaces:

  • The space {\mathrm{Sym}^1({\bf C}^2)} of linear homogeneous polynomials {L(z,w) = az + bw} of two complex variables {z,w}.
  • The space {\mathrm{Sym}^2({\bf C}^2)} of quadratic homogeneous polynomials {Q(z,w) = cz^2 + dzw + ew^2} of two complex variables {z,w}.
  • The space {\mathrm{Sym}^3({\bf C}^2)} of cubic homogeneous polynomials {C(z,w) = fz^3 + gz^2w + hzw^2 + iw^3} of two complex variables {z,w}.
(The notation {\mathrm{Sym}^k(V)} here refers to the {k^{th}} symmetric power of a vector space {V}.) Clearly these spaces are isomorphic to {{\bf C}^2, {\bf C}^3, {\bf C}^4} respectively. Furthermore, we have a multiplication map {F : \mathrm{Sym}^1({\bf C}^2) \times \mathrm{Sym}^2({\bf C}^2) \rightarrow \mathrm{Sym}^3({\bf C}^2)}, mapping a pair {(L,Q)} of a linear polynomial {L} and a quadratic polynomial {Q} to a cubic polynomial

\displaystyle F(L,Q) := LQ.

(Right now, the domain and range of this map {F} is larger dimensional than the target of three; we will cut the dimensions down to three as the argument progresses.)

The map {F}, essentially a map from {{\bf C}^5} to {{\bf C}^4}, is clearly polynomial; it is given explicitly in coordinates as

\displaystyle  F( (a,b), (c,d,e) ) = (ac, ad + bc, ae + bd, be). \ \ \ \ \ (2)

The map {F} also enjoys two basic (and commuting) symmetries:
  • If one applies a scaling {(L,Q) \mapsto (\lambda_1 L, \lambda_2 Q)} for some non-zero complex numbers {\lambda_1, \lambda_2}, then the product {LQ} is scaled by {C \mapsto \lambda_1 \lambda_2 C}: {F( \lambda_1 L, \lambda_2 Q) = \lambda_1 \lambda_2 F(L,Q)}.
  • If one applies a change of variables {(L, Q) \mapsto (L \circ T, Q \circ T)} for some invertible linear transformation {T \in \mathrm{SL}_2({\bf C})}, then the product {LQ} is transformed by {C \mapsto C \circ T}: {F(L \circ T, Q \circ T) = F(L,Q) \circ T}.
So this map enjoys a huge amount of equivariance, basically with respect to an action of the five-dimensional group {{\bf C}^\times \times {\bf C}^\times \times \mathrm{SL}_2({\bf C})}.

The five-dimensional domain {\mathrm{Sym}^1({\bf C}^2) \times \mathrm{Sym}^2({\bf C}^2)} is of course larger than the four-dimensional range {\mathrm{Sym}^3({\bf C}^2)}, so the map {F} clearly cannot be injective. This can already be seen from the scaling symmetry, as the specific scalings

\displaystyle  (L, Q) \mapsto (\lambda L, \lambda^{-1} Q) \ \ \ \ \ (3)

for {\lambda \in {\bf C}^\times} modify the linear and quadratic polynomials {L,Q} but not their product {C = LQ}. But even if one quotients out by this symmetry (3) to cut the dimension of the domain down to four, the map {F} is still not injective for the following basic reason. A generically chosen cubic polynomial {C} will split into the product {C = L_1 L_2 L_3} of three independent linear polynomials. Then there are three pairs

\displaystyle  (L_1, L_2 L_3), (L_2, L_1 L_3), (L_3, L_1 L_2) \ \ \ \ \ (4)

which all map to the same cubic polynomial

\displaystyle  F(L_1, L_2 L_3) = F(L_2, L_1 L_3) = F(L_3, L_1 L_2) = C

under the multiplication map {F}, but are not related to each other by scaling symmetry (3). Thus, we see that even after quotienting out by the scaling symmetry (3), the multiplication map {F} is generically non-injective in a three-to-one fashion. Thus we already have achieved something resembling goal (b)!

It will be convenient to “spend” the scaling symmetry {(L, Q) \mapsto (\lambda L, \lambda^{-1} Q)} to obtain a useful normalization. If {L(z,w) = az+bw} is a linear polynomial and {Q(z,w)} is a quadratic polynomial, the resultant {\mathrm{Res}(L,Q)} can be defined by the determinant

\displaystyle  \mathrm{Res}(L,Q) = \begin{vmatrix} a & b & 0 \\ 0 & a & b \\ c & d & e \end{vmatrix} = a^2 e - abd + c b^2. \ \ \ \ \ (5)

If we have a factoring

\displaystyle  L(z,w) = a (z - \alpha w), \quad Q(z,w) = c (z - \beta_1 w)(z - \beta_2 w)

then the resultant can also be described as

\displaystyle  \mathrm{Res}(L,Q) = a^2 c (\alpha - \beta_1) (\alpha - \beta_2).

Thus the resultant measures whether the linear polynomial {L} and the quadratic polynomial {Q} share a common root. A fundamental fact about resultants is that they are {SL_2}-invariant: for any {T \in SL_2({\bf C})}, we have

\displaystyle  \mathrm{Res}(L \circ T, Q \circ T) = \mathrm{Res}(L,Q).

One way to see this is to check it first for translations {(z,w) \mapsto (z + hw, w)} (which translate the roots {\alpha,\beta_1,\beta_2} by {h} while leaving {a,c} unchanged) and for inversions {(z,w) \mapsto (w,z)} (which map {\alpha,\beta_1,\beta_2} to {1/\alpha, 1/\beta_1, 1/\beta_2} while mapping {a,c} to {-a\alpha} and {c\beta_1 \beta_2} respectively), and then noting that these transformations generate all of {SL_2({\bf C})}. They also interact very nicely with scaling:

\displaystyle  \mathrm{Res}(\lambda_1 L, \lambda_2 Q) = \lambda_1^2 \lambda_2 \mathrm{Res}(L,Q).

In particular, the scaling symmetry (3) multiplies {\mathrm{Res}(L,Q)} by {\lambda}:

\displaystyle  \mathrm{Res}(\lambda L, \lambda^{-1} Q) = \lambda \mathrm{Res}(L,Q). \ \ \ \ \ (6)

Thus, we can (generically) normalize away this scaling symmetry by imposing the condition

\displaystyle  \mathrm{Res}(L,Q) = 1. \ \ \ \ \ (7)

We now have a restricted multiplication map (which by abuse of notation we will continue to call {F}) from the four-dimensional variety

\displaystyle  \{ (L,Q) \in \mathrm{Sym}^1({\bf C}^2) \times \mathrm{Sym}^2({\bf C}^2) : \mathrm{Res}(L,Q) = 1\} \ \ \ \ \ (8)

to the four-dimensional space {\mathrm{Sym}^3({\bf C}^2)}. This map {F} is still not globally injective, as we can take the three pairs in (4) from before and apply the scaling (3) separately to each of the three pairs to obtain the normalization (7). So we have kept property (b). Furthermore, this map retains the {SL_2}-equivariance (and also one remaining scaling symmetry, though we will not make much further use of that symmetry).

But we now also have property (a)! Suppose we want to show the local injectivity of {F} in the neighborhood of a pair {(L,Q)} with {\mathrm{Res}(L,Q) = 1}. As the resultant is non-vanishing, the root {\alpha} of {L} (which exists in the Riemann sphere, or projective line if you prefer) is distinct from the two roots {\beta_1, \beta_2} of {Q} (though the latter two roots could be equal to each other). Applying the {SL_2} action (which performs Möbius transforms on the roots), one can assume without loss of generality that {\alpha} is the point at infinity (or equivalently {a=0}), thus {L(z,w) = b w} for some complex number {b} and {Q(z,w) = c (z - \beta_1 w)(z - \beta_2 w)} for some complex numbers {c, \beta_1, \beta_2}, with the resultant condition (7) simplifies to {cb^2 = 1} (so in particular {c,b} are also non-zero). It is then clear that if one perturbs {L} and {Q} by a small amount (say, modifying each coefficient by {O(\varepsilon)}), then the root {\alpha=\infty} of {L} will perturb to something large ({\gg 1/\varepsilon}), while the roots {\beta_1,\beta_2} of {Q} stay bounded. Thus, just from knowledge of the product {F(L,Q)}, one can reconstruct which of the three roots of this cubic polynomial will be the perturbed root of {L}, and which two will be the perturbed roots of {Q}; from this and (6), (7) we can also reconstruct the leading coefficient {c} of {Q}, and this completely determines both {L} and {Q}. This establishes the local injectivity property (a). (In fact it is étale, but we will not need the machinery of étale maps here.)

Unfortunately, (the four-dimensional analogue of) condition (c) fails: the quadric hypersurface (8) is not isomorphic to the affine space {{\bf C}^4}. But we can try to get around this by passing to a three-dimensional slice. Let {V} be some three-dimensional affine plane of {\mathrm{Sym}^3({\bf C}^2)} (which we will take to avoid the origin for technical reasons), then we can restrict {F} as a map from the set

\displaystyle  \{ (L,Q) \in \mathrm{Sym}^1({\bf C}^2) \times \mathrm{Sym}^2({\bf C}^2) : \mathrm{Res}(L,Q) = 1; \ \ \ \ \ (9)

\displaystyle  F(L,Q) \in V\}

to {V}. The latter is clearly identifiable (by linear changes of coordinate) to {{\bf C}^3}. As {F} was already locally invertible, it remains locally invertible under restriction; and because generic cubic polynomials {C} had three preimages under {F} in (8), this continues to be the case after restricting to (9) (unless {V} was somehow so degenerate that it had no generic elements, but this turns out to be impossible). So we have retained properties (a) and (b). The miracle is that, with a good choice of {V}, we can also obtain (c) and obtain the desired counterexample to the Jacobian conjecture: despite appearances, the variety (9) is in fact equivalent to the affine space {{\bf C}^3} by polynomial changes of variable!

Let’s see how. The affine hyperplanes in {\mathrm{Sym}^3({\bf C}^2)} avoiding the origin are parameterized by the dual space of {\mathrm{Sym}^3({\bf C}^2)} avoiding the origin, which one can think of as the non-zero third order homogeneous differential operators {D = j \partial_z^3 + k \partial_z^2 \partial_w + l \partial_z \partial_w^2 + m \partial_w^3} in two variables. Indeed, every such operator {D} generates an affine hyperplane {\{ C \in \mathrm{Sym}^3({\bf C}^2) : D(C) = 1\}} that avoids the origin, and conversely by duality every affine hyperplane avoiding the origin arises in this form uniquely. Just as the cubic polynomials in {\mathrm{Sym}^3({\bf C}^2)} can be factored into three linear polynomials, the differential operators in the dual space {\mathrm{Sym}^3({\bf C}^2)^*} can also be factored into three linear differential operators, e.g.,

\displaystyle  D = j (\partial_z - \gamma_1 \partial_w) (\partial_z - \gamma_2 \partial_w) (\partial_z - \gamma_3 \partial_w)

in the case that {j} is non-zero. The {SL_2} action moves the roots {\gamma_1,\gamma_2,\gamma_3} around the Riemann sphere by Möbius transformations. As these transformations are {3}-transitive, the actual selection of such roots is not too important (and the scaling symmetry similarly makes the choice of leading coefficient {j} unimportant); the only thing to keep track of is whether the roots repeat. Up to the symmetries, there are in fact just three different equivalence classes of differential operator {D} (and thus of affine hyperplane {V}) to consider:
  • Operators where the three roots {\gamma_1,\gamma_2,\gamma_3} are all distinct, thus {D = D_1 D_2 D_3} for independent first-order operators {D_1,D_2,D_3}.
  • Operators where two roots coincide and one is distinct, thus {D = D_1^2 D_2} for independent first-order operators {D_1,D_2}.
  • Operators where all three roots coincide, thus {D = D_1^3} for some first-order operator {D_1}.

It turns out that the affine miracle for (9) occurs precisely in the second case, when {D} has two identical roots. I do not have a completely satisfactory geometric explanation for this miracle, but one can verify it by the following coordinate computation.

By applying the {SL_2} action, we can normalize so that {D = \frac{1}{2} \partial_z^2 \partial_w}, thus {V} is now the affine hyperplane of cubic polynomials {C(z,w) = f z^3 + g z^2 w + h z w^2 + i w^3} with {g=1}. Using (2) and (5), the variety (9) can now be described explicitly in coordinates as

\displaystyle  \{ (a,b,c,d,e) \in {\bf C}^5 : a^2 e - abd + cb^2 = 1; ad + bc = 1 \}. \ \ \ \ \ (10)

At first glance this seems to be a generic-looking variety cut out by a cubic equation and a quadratic equation – hardly a candidate to be affine! But observe that if {a} is non-zero, then the second equation {ad+bc = 1} can be solved for {d},

\displaystyle  d = \frac{1 - bc}{a} \ \ \ \ \ (11)

and the first equation {a^2 e - abd + cb^2 = 1} can be solved for {e},

\displaystyle  e = \frac{1 + abd - cb^2}{a^2}. \ \ \ \ \ (12)

Putting these two equations together, we see that as long as one removes the case {a=0}, the quintuple {(a,b,c,d,e)} is uniquely determined by {(a,b,c)} by a change of variables which is Laurent in {a} and polynomial in {b,c}. Thus we have a nice birational equivalence

\displaystyle  \{ (a,b,c,d,e) \in {\bf C}^5 : a^2 e - abd + cb^2 = 1; ad + bc = 1; a \neq 0 \}

\displaystyle  \cong \{ (a,b,c) \in {\bf C}^3 : a \neq 0 \}.

Thus we have already almost established property (c): the variety (9) becomes birationally equivalent to {{\bf C}^3} after cutting out the {a=0} subvariety. In particular, for each fixed non-zero value {a_0} of {a}, the corresponding fiber

\displaystyle  \{ (a,b,c,d,e) \in {\bf C}^5 : a^2 e - abd + cb^2 = 1; ad + bc = 1; a = a_0 \}

of (10) is equivalent to {{\bf C}^2} by polynomial changes of variable, since we can reconstruct {d,e} from the coordinates {b,c} by the polynomial formulae

\displaystyle  d = \frac{1-bc}{a_0}; \quad e = \frac{1 + a_0 d b - c b^2}{a_0^2}.

So we just need to glue back in the {a=0} fiber. Indeed, from (10) we see that the fiber at {0} is just

\displaystyle  \{ (0,b,c,d,e) \in {\bf C}^5 : cb^2 = 1; bc = 1 \}.

Now we observe a key miracle: the cubic equation {cb^2 = 1} and quadratic equation {bc = 1} have a unique affine solution {b=c=1} (as opposed to the six possible solutions that Bezout’s theorem might suggest – the other five solutions live on the line at infinity). So the fiber here is also affine:

\displaystyle  \{ (0,1,1,d,e) \in {\bf C}^5 : d, e \in {\bf C} \}.

This is extremely encouraging for the purposes of establishing property (c), as it strongly suggests that the variety (10) has the structure of an {{\bf C}^2}-bundle over {{\bf C}^1}, which is already extremely close to being isomorphic to the affine space {{\bf C}^3}. The main remaining task is to make sure that nothing singular happens in the limit {a \rightarrow 0}, and that a global polynomial coordinate chart for (10) that covers both the {a \neq 0} and {a = 0} fibers can be constructed.

The standard way to proceed here is to manipulate various tangent spaces using the modern machinery of algebraic geometry and commutative algebra, but given my own background, I prefer to adopt the language of analysis, and in particular big-O notation (in place of the ideals used in algebraic geometry), in order to investigate the limit {a \rightarrow 0} by hand. On the variety (10), let us use {O(X)} to denote any multiple of {X} by a polynomial expression in {a,b,c,d,e}. Thus, for instance, the equation {ad + bc = 1} implies that

\displaystyle  bc = 1 + O(a) \ \ \ \ \ (13)

while the equation {a^2 e - abd + cb^2 = 1} implies that

\displaystyle  cb^2 = 1 + O(a) \ \ \ \ \ (14)

as well as the more refined estimate

\displaystyle  cb^2 = 1 + abd + O(a^2). \ \ \ \ \ (15)

In the {a=0} case we could conclude that {b=c=1}. Now we perturb this observation. Multiplying (13) by {b} we have {b^2 c = b + O(a)}, which on substitution into (14) gives {b = 1 + O(a)}; substituting this back into either (13) or (14) also gives {c = 1 + O(a)}.

We can get some more precise asymptotics by also taking advantage of (15). Substituting {ad+bc=1} into (15), we obtain after some algebra

\displaystyle  2 cb^2 = 1 + b + O(a^2).

So if we write {b = 1+O(a)} more explicitly as {b = 1 + a y}, then we have

\displaystyle  2 c (1 + 2ay + O(a^2)) = 2 + ay + O(a^2)

and thus

\displaystyle  c = 1 - \frac{3}{2} ay + O(a^2). \ \ \ \ \ (16)

Substituting this back into (11) gives an asymptotic for {d}:

\displaystyle  d = \frac{1 - bc}{a}

\displaystyle = \frac{1 - (1 + ay) (1 - \frac{3}{2} ay + O(a^2))}{a}

\displaystyle  =\frac{1}{2} y + O(a).

Finally, one can insert these estimates into (12), although one only gets a trivial bound in this case:

\displaystyle  e = \frac{1 + abd - cb^2}{a^2}

\displaystyle = \frac{1 + a (1+O(a)) (\frac{1}{2} y + O(a)) - (1 - \frac{3}{2} ay + O(a^2)) (1 + ay)^2}{a^2}

\displaystyle  = O(1).

Expanding the {O(a^2)} error term in (16) as {a^2 z}, and doing a little more algebra, we thus have a polynomial change of variables

\displaystyle  a = a

\displaystyle  b = 1 + ay

\displaystyle  c = 1 - \frac{3}{2} ay + a^2 z

\displaystyle  d = \frac{1-bc}{a} = \frac{1}{2} y - az + \frac{3}{2} a y^2 - a^2 yz

\displaystyle  e = \frac{1 + abd - cb^2}{a^2} = -2z + 4y^2 - 4ayz + 3ay^3 - 2a^2 y^2 z

which completely parameterizes the variety (10) by polynomial combinations of three coordinates {a,y,z}. This already gives (a) and thus completes the proof of Theorem 3.

The previous computations, when expanded out, also gives polynomial inverse maps:

\displaystyle  a = a

\displaystyle  y = 2bd - ae

\displaystyle  z = 2d^2 + ce + 6bd^2 + 3bce - \frac{9}{2} e

The map from {(a,y,z)} to the {(f,h,i)} coefficients of {F(L,Q)} (dropping the {g} coefficient which is constrained to equal {1}), we obtain a polynomial map

\displaystyle  (a,y,z) \mapsto (G_1(a,y,z), G_2(a,y,z), G_3(a,y,z))

with

\displaystyle  \begin{array}{rl}  G_1(a,y,z) &= a - \frac{3}{2} a^2 y + a^3 z \\ G_2(a,y,z) &= \frac{1}{2} y - 3az + 6ay^2 - 6a^2 yz + \frac{9}{2} a^2 y^3 - 3a^3 y^2 z \\ G_3(a,y,z) &= -2z + 4y^2 - 6ayz + 7ay^3 - 6a^2 y^2 z + 3a^2 y^4 \\ & \quad - 2 a^3 y^3 z \end{array}

which theory predicts to have a constant Jacobian, and indeed one can calculate that the Jacobian is {-1}. This is essentially the original example up to trivial changes of variable; indeed, one can check that the map

\displaystyle  (a, y, -2z) \mapsto (G_3(a,y,z), 2G_2(a,y,z), 2G_1(a,y,z))

is exactly the map {F} given in (1).

AI disclosure: I used an AI chatbot to discuss various aspects of this problem and to confirm several of the calculations made here.

July 20, 2026

Secret Blogging SeminarThe new counterexample to the Jacobian conjecture

As many of you have probably heard already, yesterday morning, Levent Alpöge tweeted that Fable had found a counterexample to the Jacobian Conjecture. Specifically, let

a=(1+xy)3z+y2(1+xy)(4+3xy),b=y+3x(1+xy)2z+3xy2(4+3xy),c=2x3x2yx3z,\begin{align*} a&=&(1+xy)^3z+y^2(1+xy)(4+3xy),\\ b&=&y+3x(1+xy)^2z+3xy^2(4+3xy),\\ c&=&2x-3x^2y-x^3z, \end{align*}

Then the Jacobian of (a,b,c) is easily checked to be -2. However, the map (a,b,c) is generically three to one, not bijective.

I’m sure many of you are playing with these polynomials to see what you can figure out about them. This is a place for us to share our observations. I’ll post a few minor observations of my own soon.

First, a basic but intriguing observation from Mathoverflow user “dorky”: The polynomials a, b and c are homogeneous with respect to the grading where \deg(x) = -1, \deg(y) = 1 and \deg(z)=2; their degrees are \deg(a) = 2, \deg(b) = 1 and \deg(c) = -1. I’m not sure what to make of this, but it surely matters.


Some computations by me: If you eliminate any two of the variables (x,y,z), you get a cubic relation in the remaining variable. Here they are

2c+(43bc)x+(16ab218abc+b3c+27a2c2)x3(18ab+b3+27a2c)+18ay3by2+2y3(really long)+8z3\begin{matrix} -2 c+(4 – 3 b c) x + (16 a – b^2 – 18 a b c + b^3 c + 27 a^2 c^2) x^3 \\ (-18 a b + b^3 + 27 a^2 c)+18 ay-3 b y^2+ 2y^3 \\ (\text{really long}) + 8 z^3 \\ \end{matrix}

I’m leaving out the “really long”, because it is really long and I suspect we don’t care about the details. Put

Δ=16ab218abc+b3c+27a2c2\Delta= 16 a – b^2 – 18 a b c + b^3 c + 27 a^2 c^2 ,

the leading coefficient of the x cubic. Then the discriminants of the three cubics are \Delta p^2, \Delta q^2, \Delta r^2 where

pamp;=amp;89bc+27ac2qamp;=amp;bramp;=amp;(really long)\begin{align*} p &=& 8 – 9 b c + 27 a c^2 \\ q &=& b \\ r &=& (\text{really long}) \\ \end{align*}

The polynomials (p,q,r) have no common zeroes. Roughly speaking, our map should have special behavior over the loci \Delta=0, p=0, q=0 and r=0. The fact that $p$, $q$ and $r$ each appear cubed means that the variables x, y and z should have three fold branching over the loci p=0, q=0 and r=0 (respectively).

I’m having trouble visualizing what happens over \Delta=0 — since the leading coefficient of the x cubic drops out, the map is 2 to 1 rather than 3 to 1 over this point. But, at the same time, the y and z cubics have a multiple root at the points of \Delta=0. Does anyone see how to visualize this?

Any other insights?

Tommaso DorigoToward Mode Collapse of Natural Language

Toward Mode Collapse of Natural Language

Regression toward the mean is a simple phenomenon commonly described in Statistics 101 courses. If you measure a parameter describing some phenomenon, you will find that extreme measured values tend to be followed by less extreme ones.

Tommaso Dorigo
Categories

July 19, 2026

Jordan EllenbergProposed World Cup rules change

Correct outcome, but imagine if Argentina had won this on penalties without taking a single shot on goal in 120 minutes of actual play! That could easily have happened, and what an embarrassment.

My proposed rule change: the number of penalty kicks a team gets is either five or the number of shots taken on the field of play, whichever is fewer. Couldn’t mount a serious attack the whole game and only made two shots on goal? Too bad, so sad.

John BaezGalilean Limits of Electromagnetism

Maxwell’s equations are invariant under Lorentz transformations. The usual equations of fluid flow are not! Like the rest of Newtonian mechanics, they’re invariant under Galilean transformations like

t' = t,  \quad  x' = x - vt

So, if we simply slap these two theories together, we get a mess! How can we study electrically conductive fluids—like plasma—without bringing special relativity into the game?

We can use a limiting case of Maxwell’s equations where we ignore terms that become tiny when all the particles are moving much slower than light.

There seem to be at least two ways to do this: there’s an ‘electric limit’ of Maxwell’s equations and a ‘magnetic limit’. Both are invariant under Galilean transformations. The original derivation of these limits by Le Bellac and Lévy-Leblond in 1973 used the version of Maxwell’s equations including the electric permittivity \varepsilon_0 and magnetic permeability \mu_0 of the vacuum, whose product is 1/c^2. This is convenient but not necessary, as explained here:

• Jose A. Heras, The Galilean limits of Maxwell’s equations.

In the magnetic limit of Maxwell’s equations, we throw out effects due to time-varying electric fields:



People often use the magnetic limit when studying nonrelativistic electrically conductive fluids. In this situation they often consider a version of the magnetic limit where the charge density \rho is zero, since this is typically close to true in a plasma. However Heras does not do this, nor does the original paper:

• Le Bellac and Levy-Leblond, Galilean electromagnetism.

In the electric limit of Maxwell’s equations, we throw out effects due to time-varying magnetic fields:



It’s fun to compare the magnetic and electric limits.

The magnetic limit has been called ‘pre-Maxwellian’, because it’s like electromagnetism before Maxwell added the extra term that makes a changing electric field create a curl in the magnetic field. Without this term there is no light!

In the electric limit you also can’t have light, because it’s missing the term that makes a changing magnetic field create a curl in the electric field.

In the magnetic limit you can’t have capacitors, because those store energy in the electric field, and in the magnetic limit the energy density is just \mathbf{B} \cdot \mathbf{B}/2.

Similarly, in the electric limit you can’t have inductors, because inductors store energy in the magnetic field, and in this limit the energy density is just \mathbf{E} \cdot \mathbf{E}/2.

It’s all nicely symmetrical! But still somewhat mysterious to me. All the derivations of these limits that I’ve seen involve too many parameters for my taste, and too much talk. But that’s how I often feel when I’m just starting to study a piece of physics.

Besides the two papers mentioned in my last post, I’ve been looking at this:

• Giovanni Manfredi, Non-relativistic limits of Maxwell’s equations.

There’s a lot I haven’t explained here. I haven’t even said how the electric or magnetic fields transform under Galilean boosts in these limiting theories! I find this subject fairly confusing, and I’d probably have to redo all the calculations to really understand them. As Feynman said, “what I cannot create I do not understand”.

Someday I should dig deeper into this subject and explain how the two limits work in a way I find satisfying. I should also draw the connections to this earlier article of mine:

Magnetohydrodynamics.

John BaezJordan Triples and the Standard Model

I don’t usually talk about particle physics here. I have a whole series of articles about octonions and the Standard Model on my other blog. But I’m kind of excited about this new paper, so I’ll talk about it here too:

• John Baez, Endre Bokor and Latham Boyle, Jordan pair quantum theory and the Standard Model.

Jordan algebras were introduced by Jordan, von Neumann and Wigner in 1934 in an attempt to formalize algebras of observables in quantum theory. They come in 4 infinite series—but there’s one more, the ‘exceptional Jordan algebra’, consisting of 3 × 3 self-adjoint matrices of octonions. For years physicists sought to find some use for it.

In 2018, Todorov and Dubois–Violette noticed that the symmetries of the exceptional Jordan include the Standard Model gauge group in a nice way. But it was unclear how to bring in the fermions—the quarks and leptons. That’s what our new paper does.

To do this, we need to go beyond Jordan algebras. Jordan pairs and Jordan triples are two closely linked formalisms that generalize Jordan algebras. Our paper explains them in detail—and how they’re connected to geometry and quantum mechanics. But here I will mostly skip that wonderful story, so I can quickly explain the connection to the Standard Model.

Here’s how the Standard Model gauge group, together with its representation on one generation of fermions, drops out of a Jordan triple.

The bi-Cayley triple

Let

\mathbb{O}_\mathbb{C} = \mathbb{C} \textstyle{\otimes}_\mathbb{R} \mathbb{O}

be the bioctonions: octonions with complex coefficients. Write \mathbb{O}_\mathbb{C}^2 for the space of column vectors with two bioctonion entries.

\mathbb{O}_\mathbb{C}^2 has a certain triple product

[x,y,z]=\frac{1}{2}(x(y^{\dagger}z)+z(y^{\dagger}x))

which obey the axioms of a gadget called a ‘positive hermitian Jordan triple’. It’s called the bi-Cayley triple.

Now, every positive hermitian Jordan triple gives rise to a \mathbb{Z}_2-graded real Lie algebra

\mathbf{k} = \mathbf{k}_0 \textstyle{\oplus} \mathbf{k}_1

Not a Lie superalgebra: a plain old-fashioned Lie algebra with a \mathbb{Z}_2-grading!

How does this work? We take the hermitian Jordan triple itself to be \mathbf{k}_1. The Lie algebra \mathbf{k}_0 consists of all linear maps from \mathbf{k}_1 to itself that are of this form:

x \mapsto [a,b,x] - [b,a,x]

for some a,b \in \mathbf{k}_1. These maps are called real inner derivations. They form a Lie algebra since the commutator of two such maps is another such map. With a bit more work we can define other operations making all of \mathbf{k} into a \mathbb{Z}_2-graded Lie algebra.

So, we get a big Lie algebra \mathbf{k}, and a Lie subalgebra \mathbf{k}_0 sitting inside it. From this we get two Lie groups: a big one K whose Lie algebra is \mathbf{k}, and a subgroup K_0 whose Lie algebra is \mathbf{k}_0.

The quotient is K/K_0 is a nice kind of manifold called a hermitian symmetric space. Conversely, any compact hermitian symmetric space give rise to a positive hermitian Jordan triple!

This geometric picture is revealing. The group K acts transitively as symmetries of our hermitian symmetric space, while the stabilizer of any point is isomorphic to K_0. Our original Jordan triple, \mathbf{k}_1, is then the tangent space of that point! So, K_0 acts on this Jordan triple. This action preserves the triple product, and we call K_0 the real inner automorphism group of our Jordan triple.

Here’s another great thing about the geometric picture: hermitian symmetric spaces were classified by Eli Cartan (who seems to have spent his life classifying things). As a result we also know the classification of positive hermitian Jordan triples. They come in four infinite series together with two exceptions. One is the bi-Cayley triple, and other is the Albert triple, which is the complexification of the exceptional Jordan algebra. The bi-Cayley triple is a subtriple of the Albert triple. It’s these two exceptions that are connected to the Standard Model. But we’ll start with the bi-Cayley triple.

The 3-graded Lie algebra coming from the bi-Cayley triple is the compact real form of \mathfrak{e}_6:

\mathfrak{e}_6 = \big[\mathfrak{so}(10) \textstyle{\oplus} \mathfrak{u}(1)\big] \textstyle{\oplus} \mathbb{O}_\mathbb{C}^2

The even part of this Lie algebra is in brackets. The corresponding hermitian symmetric space is called the bioctonionic plane (\mathbb{C}\otimes\mathbb{O})P^2. The even part of our 3-graded Lie algebra, \mathfrak{so}(10)\oplus \mathfrak{u}(1), generates the stabilizer of a point in the bioctonionic plane. The odd part, our friend \mathbb{O}_\mathbb{C}^2, is the tangent space of that point.

Here’s the first big surprise. The even part transforms as the adjoint representation of \mathrm{Spin}(10), while the odd part itself transforms as the 16-dimensional complex spinor representation of \mathrm{Spin}(10). Ignoring the extra \mathrm{U}(1) for a moment, this is exactly what we see in a \mathrm{SO}(10) grand unified theory: gauge bosons in the adjoint representation, and one generation of fermions in the 16-dimensional spinor representation.

So before we do anything, the bi-Cayley triple already smells like it contains the ingredients of an \mathrm{SO}(10) grand unified theory.

Tripotents

In a Jordan algebra the important elements are the idempotents, e^2 = e. In a Jordan triple W their role is played by tripotents: elements e with

[e,e,e] = e

A tripotent always lets us split W into three parts via something called its Peirce decomposition. The operator w \mapsto [e,e,w] has eigenvalues 0, 1/2, and 1, so W splits into the corresponding eigenspaces

W = W_0(e) \textstyle{\oplus} W_{1/2}(e) \textstyle{\oplus} W_1(e)

which are called the Peirce 0-space, Peirce 1/2-space and Peirce 1-space of e. A tripotent is called minimal when its Peirce 1-space is one-dimensional. Two tripotents e_1, e_2 are called colinear when each lies in the other’s Peirce 1/2-space.

I can’t resist explaining some of the quantum physics here. In a hermitian Jordan triple, the triple product [-,-,-] is linear in the first and last slot, but conjugate-linear in the middle slot. So, if you multiply a tripotent by a phase \alpha, you get a new tripotent:

[\alpha e, \alpha e, \alpha e] = \alpha \overline{\alpha} \alpha e = \alpha e

This should remind you of how when you multiply a unit vector in a Hilbert space by a phase, you get a new unit vector. In Jordan triple quantum mechanics, minimal tripotents take the place of these unit vectors. The hermitian symmetric space K/K_0 that I was talking about earlier is the same as the space of minimal tripotents mod phase! So, it generalizes the familiar space of ‘pure states’ in quantum mechanics: unit vectors mod phase.

But let’s get back to the Standard Model.

A chain of Jordan triples

From here on, the single fact driving everything is this: in any hermitian Jordan triple, any minimal tripotent’s Peirce 1/2-space is itself a hermitian Jordan triple!

If we run this starting from the bi-Cayley triple, we get this chain of hermitian Jordan triples, where each row’s 1/2-space is the next row’s triple:

Jordan triple Lie algebra \mathbf{k}_0 \oplus \mathbf{k}_1 (even part in brackets)
W = \mathbb{O}_\mathbb{C}^2 \mathfrak{e}_6 = [\mathfrak{so}(10) \oplus \mathfrak{u}(1)] \oplus \mathbb{O}_\mathbb{C}^2
W' = \mathfrak{a}_5(\mathbb{C}) \mathfrak{so}(10) = [\mathfrak{su}(5) \oplus \mathfrak{u}(1)] \oplus \mathfrak{a}_5(\mathbb{C})
W'' = \mathrm{M}_{3,2}(\mathbb{C}) \mathfrak{su}(5) = [\mathfrak{g}_{\mathrm{SM}}] \oplus \mathrm{M}_{3,2}(\mathbb{C})

Here \mathfrak{a}_5(\mathbb{C}) is the Jordan triple of antisymmetric 5\times 5 complex matrices, \mathrm{M}_{3,2}(\mathbb{C}) is the Jordan triple of 3\times 2 complex matrices, \mathfrak{g}_{\mathrm{SM}} = \mathfrak{su}(3)\oplus\mathfrak{su}(2)\oplus \mathfrak{u}(1), and

G_{\mathrm{SM}} = \mathrm{S}(\mathrm{U}(2) \times \mathrm{U}(3)) \cong (\mathrm{SU}(3)\times\mathrm{SU}(2)\times\mathrm{U}(1))/\mathbb{Z}_6

is the true Standard Model gauge group.

The gauge group from two tripotents

Start with the bi-Cayley triple. Choose two colinear minimal tripotents e_1, e_2. Descend the table twice:

• Start with W = \mathbb{O}_\mathbb{C}^2, which has real inner automorphism group (\mathrm{Spin}(10)\times\mathrm{U}(1))/\mathbb{Z}_4.

• Fix e_1. Its Peirce 1/2-space is latex W’ = \mathfrak{a}_5(\mathbb{C}),$ with real inner automorphism group \mathrm{SU}(5)\times\mathrm{U}(1).

• Fix e_2 (colinear with e_1, so living in W'). Its Peirce 1/2-space in latex W’$ is W'' = \mathrm{M}_{3,2}(\mathbb{C}), with real inner automorphism group exactly G_{\mathrm{SM}}.

In other words, the subspace of the bi-Cayley triple colinear with both e_1 and e_2 is a Jordan triple whose real inner automorphism group is the Standard Model gauge group.

The choice of e_1 and e_2 also pins down how G_{\mathrm{SM}} sits inside the original group \mathrm{E}_6. At each we step take the subgroup that acts with determinant 1 and preserves the chosen tripotent up to a phase; this gives a chain of subgroups whose members are \mathrm{Spin}(10), \mathrm{U}(5), and G_{\mathrm{SM}}, so we get the embeddings

G_{\mathrm{SM}} \subset \mathrm{SU}(5) \subset \mathrm{Spin}(10)

In particle physics, this is the classic chain taking us from the so-called \mathrm{SO}(10) grand unified theory down to the \mathrm{SU}(5) grand unified theory down to the Standard Model. And it’s well known that restricting the 16-dimensional complex spinor representation of \mathrm{Spin}(10) along this chain gives precisely the Standard Model representation \rho_{\mathrm{SM}} on one generation of fermions! So we get one generation of Standard Model fermions this way.

The six particles types as Peirce spaces

We have gotten the representation of the Standard Model gauge group on one generation of fermions without any fuss. But it’s also fun to peer into the details, and see how the different kinds of fermions emerge.

For any tripotent e, we have projections P_0(e), P_{1/2}(e) and P_1(e) onto its three eigenspaces: its so-called Peirce projectors. Since we get the Standard Model structure using two minimal tripotents e_1 and e_2 in the bi-Cayley triple \mathbb{O}_{\mathbb{C}}^2, there are nine composites of two Peirce projectors we can apply to this triple. This is how we pick out the different kinds of fermions!

As a representation of the Standard Model Lie algebra

\mathfrak{g}_{\mathrm{SM}} = \mathfrak{su}(3) \textstyle{\oplus} \mathfrak{su}(2) \textstyle{\oplus} \mathfrak{u}(1)

any generation of Standard Model fermions transforms as the direct sum of six irreducible representations:

\rho_{\mathrm{SM}} = (3,2,\tfrac{1}{6}) \textstyle{\oplus} (\bar 3,1,\tfrac{1}{3}) \textstyle{\oplus} (\bar 3,1,-\tfrac{2}{3}) \textstyle{\oplus} (1,2,-\tfrac{1}{2}) \textstyle{\oplus} (1,1,1) \textstyle{\oplus} (1,1,0)

These correspond to the six types of left-handed fermion: q_L, \overline{d_R}, \overline{u_R}, \ell_L, \overline{e_R}, \overline{\nu_R}. Six irreducible pieces, six particle types.

It turns out these are exactly the six nonzero components of the Peirce decomposition of \mathbb{O}_\mathbb{C}^2 with respect to both e_1 and e_2. Those six match up one-to-one with the particle types:

Peirce projector representation of G_{\text{SM}} particle type
P_{1/2}(e_2) P_{1/2}(e_1) (3, 2, +1/6) q_L
P_{1/2}(e_2) P_0(e_1) (\overline{3}, 1, +1/3) \overline{d_R}
P_0(e_2) P_{1/2}(e_1) (\overline{3}, 1, −2/3) \overline{u_R}
P_0(e_2) P_0(e_1) (1, 2, −1/2) \ell_L
P_1(e_2) P_{1/2}(e_1) (1, 1, +1) \overline{e_R}
P_{1/2}(e_2) P_1(e_1) (1, 1, 0) \overline{\nu_R}

The remaining three combinations—P_1(e_2)P_1(e_1), P_1(e_2)P_0(e_1), and P_0(e_2)P_1(e_1)—all vanish, which is why we land on six pieces and not nine.

So the whole package—the gauge group G_{\mathrm{SM}}, the embedding G_{\mathrm{SM}} \subset \mathrm{Spin}(10), the representation \rho_{\mathrm{SM}}, and even the split of one generation into its six particle multiplets as distinct Peirce components—all comes out of the single object \mathbb{O}_\mathbb{C}^2 once you choose two colinear minimal tripotents.

And if you prefer to start one level up, with the Albert triple \mathfrak{h}_3(\mathbb{O}) \otimes \mathbb{C}, you get the same result by choosing three mutually colinear tripotents instead of two—but for that, read our paper!

n-Category Café Octonions and the Standard Model (Part 15)

Last time I described a way to get the Standard Model gauge group from the exceptional Jordan algebra. But that approach gave no obvious nice way to put quarks and leptons into the picture. This new paper tackles that problem:

Jordan pairs and Jordan triples are two closely linked formalisms that generalize Jordan algebras. Our paper explains them in detail — and how they’re connected to geometry and quantum mechanics. Here I will mostly skip that wonderful story, so I can quickly explain the connection to the Standard Model.

Here’s how the Standard Model gauge group, together with its representation on one generation of fermions, drops out of a Jordan triple.

The bi-Cayley triple

Let

𝕆 = 𝕆\mathbb{O}_\mathbb{C} = \mathbb{C} \textstyle{\otimes}_\mathbb{R} \mathbb{O}

be the bioctonions: octonions with complex coefficients. Write 𝕆 2\mathbb{O}_\mathbb{C}^2 for the space of column vectors with two bioctonion entries.

𝕆 2\mathbb{O}_\mathbb{C}^2 has a certain triple product

[x,y,z]=12(x(y z)+z(y x)) [x,y,z]=\frac{1}{2}(x(y^{\dagger}z)+z(y^{\dagger}x))

which obey the axioms of a gadget called a ‘positive hermitian Jordan triple’. It’s called the bi-Cayley triple.

Now, every positive hermitian Jordan triple gives rise to a 2\mathbb{Z}_2-graded real Lie algebra

k=k 0k 1 \mathbf{k} = \mathbf{k}_0 \textstyle{\oplus} \mathbf{k}_1

Not a Lie superalgebra: a plain old-fashioned Lie algebra with a 2\mathbb{Z}_2-grading!

How does this work? We take the hermitian Jordan triple itself to be k 1\mathbf{k}_1. The Lie algebra k 0\mathbf{k}_0 consists of all linear maps from k 1\mathbf{k}_1 to itself that are of this form:

x[a,b,x][b,a,x] x \mapsto [a,b,x] - [b,a,x]

for some a,bk 1a,b \in \mathbf{k}_1. These maps are called real inner derivations. They form a Lie algebra since the commutator of two such maps is another such map. With a bit more work we can define other operations making all of k\mathbf{k} into a 2\mathbb{Z}_2-graded Lie algebra.

So, we get a big Lie algebra k\mathbf{k}, and a Lie subalgebra k 0\mathbf{k}_0 sitting inside it. From this we get two Lie groups: a big one KK whose Lie algebra is k\mathbf{k}, and a subgroup K 0K_0, whose Lie algebra is k 0\mathbf{k}_0.

The quotient is K/K 0K/K_0 is a nice kind of manifold called a hermitian symmetric space. Conversely, any compact hermitian symmetric space give rise to a positive hermitian Jordan triple!

This geometric picture is revealing. The group KK acts transitively as symmetries of our hermitian symmetric space, while the stabilizer of any point is isomorphic to K 0K_0. Our original Jordan triple, k 1\mathbf{k}_1, is then the tangent space of that point. So, K 0K_0 acts on the Jordan triple. This action preserves the triple product, and we call K 0K_0 the real inner automorphism group of our Jordan triple.

Here’s another great thing about the geometric picture: hermitian symmetric spaces were classified by Eli Cartan (who seems to have spent his life classifying things). As a result we also know the classification of positive hermitian Jordan triples. They come in four infinite series together with two exceptions. One is the bi-Cayley triple, and other is the Albert triple, which is the complexification of the exceptional Jordan algebra. The bi-Cayley triple is a subtriple of the Albert triple. It’s these two exceptions that are connected to the Standard Model. But we’ll start with the bi-Cayley triple.

The 3-graded Lie algebra coming from the bi-Cayley triple is the compact real form of 𝔢 6\mathfrak{e}_6:

𝔢 6=[𝔰𝔬(10)𝔲(1)]𝕆 2.\mathfrak{e}_6 = \big[\mathfrak{so}(10) \textstyle{\oplus} \mathfrak{u}(1)\big] \textstyle{\oplus} \mathbb{O}_\mathbb{C}^2.

The even part of this Lie algebra is in brackets. The corresponding hermitian symmetric space is called the bioctonionic plane (𝕆)P 2(\mathbb{C}\otimes\mathbb{O})P^2. I explained it in Part 12. The even part of our 3-graded Lie algebra, 𝔰𝔬(10)𝔲(1)\mathfrak{so}(10)\oplus \mathfrak{u}(1), generates the stabilizer of a point in the bioctonionic plane. The odd part, our friend 𝕆 2\mathbb{O}_\mathbb{C}^2, is the tangent space of that point.

Here’s the first big surprise. The even part transforms as the adjoint representation of Spin(10)\mathrm{Spin}(10), while the odd part itself transforms as the 16-dimensional complex spinor representation of Spin(10)\mathrm{Spin}(10). Ignoring the extra U(1)\mathrm{U}(1) for a moment, this is exactly what we see in a SO(10)\mathrm{SO}(10) grand unified theory: gauge bosons in the adjoint representation, and one generation of fermions in the 16-dimensional spinor representation.

So before we do anything, the bi-Cayley triple already smells like it contains the ingredients of an SO(10)\mathrm{SO}(10) grand unified theory.

Tripotents

In a Jordan algebra the important elements are the idempotents, e 2=ee^2 = e. In a Jordan triple WW their role is played by tripotents: elements ee with

[e,e,e]=e.[e,e,e] = e.

A tripotent always lets us split WW into three parts via something called its Peirce decomposition. The operator w[e,e,w]w \mapsto [e,e,w] has eigenvalues 0,12,10, \tfrac{1}{2}, 1, and WW splits into the corresponding eigenspaces

W=W 0(e)W 1/2(e)W 1(e),W = W_0(e) \textstyle{\oplus} W_{1/2}(e) \textstyle{\oplus} W_1(e),

which are called the Peirce 0-space, Peirce 12\tfrac{1}{2}-space and Peirce 1-space of ee. A tripotent is called minimal when its Peirce 11-space is one-dimensional: minimal tripotents are the analogues of unit vectors in ordinary quantum theory. Two tripotents e 1,e 2e_1, e_2 are called colinear when each lies in the other’s Peirce 12\tfrac{1}{2}-space.

I can’t resist explaining some of the quantum physics here. I said I wouldn’t, but I can’t help it. In a hermitian Jordan triple, the triple product [,,][-,-,-] is linear in the first and last slot, but conjugate-linear in the middle slot. So, if you multiply a tripotent by a phase α\alpha, you get a new tripotent:

[αe,αe,αe]=αα¯αe=αe [\alpha e, \alpha e, \alpha e] = \alpha \overline{\alpha} \alpha e = \alpha e

This should remind you of how when you multiply a unit vector in a Hilbert space by a phase, you get a new unit vector. In Jordan triple quantum mechanics, minimal tripotents take the place of these unit vectors. And guess what: the hermitian symmetric space K/K 0K/K_0 that I was talking about earlier is also the space of minimal tripotents mod phase! So, it generalizes the familiar space of ‘pure states’ in quantum mechanics, which are unit vectors mod phase.

But let’s get back to the Standard Model.

A chain of Jordan triples

From here on, the single fact driving everything is this: in any hermitian Jordan triple, any minimal tripotent’s Peirce 12\tfrac{1}{2}-space is itself a hermitian Jordan triple!

If we run this starting from the bi-Cayley triple, we get this chain of hermitian Jordan triples, where each row’s 12\tfrac{1}{2}-space is the next row’s triple:

Jordan triple Lie algebra k 0k 1\mathbf{k}_0 \oplus \mathbf{k}_1 (even part in brackets) real inner automorphism group
W=𝕆 2W = \mathbb{O}_\mathbb{C}^2 𝔢 6=[𝔰𝔬(10)𝔲(1)]𝕆 2\mathfrak{e}_6 = [\mathfrak{so}(10) \oplus \mathfrak{u}(1)] \oplus \mathbb{O}_\mathbb{C}^2 (Spin(10)×U(1))/ 4(\mathrm{Spin}(10) \times \mathrm{U}(1)) / \mathbb{Z}_4
W=𝔞 5()W' = \mathfrak{a}_5(\mathbb{C}) 𝔰𝔬(10)=[𝔰𝔲(5)𝔲(1)]𝔞 5()\mathfrak{so}(10) = [\mathfrak{su}(5) \oplus \mathfrak{u}(1)] \oplus \mathfrak{a}_5(\mathbb{C}) SU(5)×U(1)\mathrm{SU}(5) \times \mathrm{U}(1)
W=M 3,2()W'' = \mathrm{M}_{3,2}(\mathbb{C}) 𝔰𝔲(5)=[𝔤 SM]M 3,2()\mathfrak{su}(5) = [\mathfrak{g}_{\mathrm{SM}}] \oplus \mathrm{M}_{3,2}(\mathbb{C}) G SMG_{\mathrm{SM}}

Here 𝔞 5()\mathfrak{a}_5(\mathbb{C}) is the Jordan triple of antisymmetric 5×55\times 5 complex matrices, M 3,2()\mathrm{M}_{3,2}(\mathbb{C}) is the Jordan triple of 3×23\times 2 complex matrices, 𝔤 SM=𝔰𝔲(3)𝔰𝔲(2)𝔲(1)\mathfrak{g}_{\mathrm{SM}} = \mathfrak{su}(3)\oplus\mathfrak{su}(2)\oplus \mathfrak{u}(1), and

G SM=S(U(2)×U(3))(SU(3)×SU(2)×U(1))/ 6G_{\mathrm{SM}} = \mathrm{S}(\mathrm{U}(2) \times \mathrm{U}(3)) \cong (\mathrm{SU}(3)\times\mathrm{SU}(2)\times\mathrm{U}(1))/\mathbb{Z}_6

is the true Standard Model gauge group.

The gauge group from two tripotents

Now pick two colinear minimal tripotents e 1,e 2We_1, e_2 \in W. Descend the table twice:

  • Start with W=𝕆 2W = \mathbb{O}_\mathbb{C}^2, which has real inner automorphism group (Spin(10)×U(1))/ 4(\mathrm{Spin}(10)\times\mathrm{U}(1))/\mathbb{Z}_4.
  • Fix e 1e_1. Its Peirce 12\tfrac{1}{2}-space is W=𝔞 5()W' = \mathfrak{a}_5(\mathbb{C}), with real inner automorphism group SU(5)×U(1)\mathrm{SU}(5)\times\mathrm{U}(1).
  • Fix e 2e_2 (colinear with e 1e_1, so living in WW'). Its Peirce 12\tfrac{1}{2}-space in WW' is W=M 3,2()W'' = \mathrm{M}_{3,2}(\mathbb{C}), with real inner automorphism group exactly G SMG_{\mathrm{SM}}.

In other words, the subspace of the bi-Cayley triple colinear with both e 1e_1 and e 2e_2 is a Jordan triple whose real inner automorphism group is the Standard Model gauge group.

The choice of e 1e_1 and e 2e_2 also pins down how G SMG_{\mathrm{SM}} sits inside the original group E 6\mathrm{E}_6. At each we step take the subgroup that acts with determinant 11 and preserves the chosen tripotent up to a phase; this gives a chain of subgroups whose members are Spin(10)\mathrm{Spin}(10), U(5)\mathrm{U}(5), and G SMG_{\mathrm{SM}}, so we get the embedding

G SMSU(5)Spin(10). G_{\mathrm{SM}} \subset \mathrm{SU}(5) \subset \mathrm{Spin}(10).

In particle physics, this is the classic chain taking us from the so-called SO(10)\mathrm{SO}(10) grand unified theory down to the SU(5)\mathrm{SU}(5) grand unified theory down to the Standard Model. And it’s well known that restricting the 16-dimensional complex spinor representation of Spin(10)\mathrm{Spin}(10) along this chain gives precisely the Standard Model representation ρ SM\rho_{\mathrm{SM}} on one generation of fermions! So we get one generation of Standard Model fermions this way.

The six particles types as Peirce spaces

We have gotten the representation of the Standard Model gauge group on one generation of fermions without any fuss. But it’s also fun to peer into the details, and see how the different kinds of fermions emerge. We can get them using the fact that for any tripotent ee, we have projections P 0(e),P 1/2(e)P_0(e), P_{1/2}(e) and P 1(e)P_1(e) onto its three eigenspaces: its so-called Peirce projectors.

Since we get the Standard Model gauge group and its representation on fermions from two minimal tripotents e 1e_1 and e 2e_2, we have nine Peirce projectors we can apply to our Jordan triple 𝕆 2\mathbb{O}_{\mathbb{C}}^2. Let’s use these to pick out various kinds of particles!

As a representation of the Standard Model Lie algebra

𝔤 SM=𝔰𝔲(3)𝔰𝔲(2)𝔲(1),\mathfrak{g}_{\mathrm{SM}} = \mathfrak{su}(3) \textstyle{\oplus} \mathfrak{su}(2) \textstyle{\oplus} \mathfrak{u}(1) ,

any generation of Standard Model fermions transforms as the direct sum of six irreducible representations:

ρ SM=(3,2,16)(3¯,1,13)(3¯,1,23)(1,2,12)(1,1,1)(1,1,0),\rho_{\mathrm{SM}} = (3,2,\tfrac{1}{6}) \textstyle{\oplus} (\bar 3,1,\tfrac{1}{3}) \textstyle{\oplus} (\bar 3,1,-\tfrac{2}{3}) \textstyle{\oplus} (1,2,-\tfrac{1}{2}) \textstyle{\oplus} (1,1,1) \textstyle{\oplus} (1,1,0),

These correspond to the six types of left-handed fermion: q L,d R¯,u R¯, L,e R¯,ν R¯q_L, \overline{d_R}, \overline{u_R}, \ell_L, \overline{e_R}, \overline{\nu_R}. Six irreducible pieces, six particle types.

It turns out these are exactly the six nonzero components of the Peirce decomposition of 𝕆 2\mathbb{O}_\mathbb{C}^2 with respect to both e 1e_1 and e 2e_2. Those six match up one-to-one with the particle types:

Peirce projector representation of G SMG_{\text{SM}} particle type
P 1/2(e 2)P 1/2(e 1)P_{1/2}(e_2) P_{1/2}(e_1) (3, 2, +1/6) q Lq_L
P 1/2(e 2)P 0(e 1)P_{1/2}(e_2) P_0(e_1) (3¯\overline{3}, 1, +1/3) d R¯\overline{d_R}
P 0(e 2)P 1/2(e 1)P_0(e_2) P_{1/2}(e_1) (3¯\overline{3}, 1, −2/3) u R¯\overline{u_R}
P 0(e 2)P 0(e 1)P_0(e_2) P_0(e_1) (1, 2, −1/2) L\ell_L
P 1(e 2)P 1/2(e 1)P_1(e_2) P_{1/2}(e_1) (1, 1, +1) e R¯\overline{e_R}
P 1/2(e 2)P 1(e 1)P_{1/2}(e_2) P_1(e_1) (1, 1, 0) ν R¯\overline{\nu_R}

The remaining three combinations — P 1(e 2)P 1(e 1)P_1(e_2)P_1(e_1), P 1(e 2)P 0(e 1)P_1(e_2)P_0(e_1), and P 0(e 2)P 1(e 1)P_0(e_2)P_1(e_1) — all vanish, which is why we land on six pieces and not nine.

So the whole package — the gauge group G SMG_{\mathrm{SM}}, the embedding G SMSpin(10)G_{\mathrm{SM}} \subset \mathrm{Spin}(10), the representation ρ SM\rho_{\mathrm{SM}}, and even the split of one generation into its six particle multiplets as distinct Peirce components — all comes out of the single object 𝕆 2\mathbb{O}_\mathbb{C}^2 once you choose two colinear minimal tripotents.

And if you prefer to start one level up, with the Albert triple 𝔥 3(𝕆)\mathfrak{h}_3(\mathbb{O}) \otimes \mathbb{C}, you get the same result by choosing three mutually colinear tripotents instead of two — but for that, read our paper!

Scott Aaronson NISQ and quantum supremacy did not fail

A week ago, a philosopher named Amit Hagar put out a preprint entitled The NISQ Trap: Eight Years of Demonstrations the Hardware was Built to Lose. Here’s the abstract:

With a single clear exception, every NISQ-era flagship demonstration of ‘quantum advantage’ has, within eighteen months of its announcement, been classically reproduced, shown to rest on classically tractable structure, or closed by a simulability theorem. Six theoretical results from 2024 through April 2026 explain the pattern: the regions of circuit-space NISQ hardware can run with sufficient fidelity coincide with the regions classical algorithms compress efficiently, because the features that admit one (low effective depth, strong algebraic structure, geometric locality) are the features that admit the other. This reading dates the NISQ programme from its 2018 articulation as an interim retreat from the unmet conditions of the 1996 threshold theorems, characterises the eight years that followed as a closed loop in which the demonstrations the hardware could run were drawn from regions classical methods could already reach, and locates the exit from the loop where the threshold theorems originally located it: in fault tolerance. The empirical pattern could in principle break with a demonstration that escapes the current simulability results. After eight years and more than thirty advantage-class announcements, the burden of producing such a demonstration falls to the defenders of NISQ.

You can also read some debates about the paper on SciRate here. I think it’s fair to say that the paper is purely polemical, without new ideas, and Pangram agrees with my suspicion (and that of a SciRate commenter) that significant portions of it are AI-generated.

Nevertheless, the basic thesis—that quantum supremacy in the NISQ (Noisy Intermediate Scale Quantum computing) era has been a failure, or even an example of pathological science—seems surprisingly widely shared, along with the opposite thesis that quantum computing already gives oodles of useful advantages for optimization and finance.

So it seems worth stating for the record that I have an extremely different view. I would say:

  1. Sampling-based quantum supremacy experiments, including those based on Random Circuit Sampling and BosonSampling, passed the point about two years ago where, absent a breakthrough in classical algorithms, they quite clearly are beating what can easily be simulated on any existing classical computer. Hagar seems to claim that these experiments have been killed by the October 2025 paper Classical simulation of noisy random circuits from exponential decay of correlation, but he ignores that the algorithm from that paper still needs time that’s exponential in the circuit depth (see Theorem 2).
  2. Indeed, simulating deep ~100-qubit random circuits, like those that Google and Quantinuum have now demonstrated experimentally, still seems pretty hopeless with any current classical method. This is particularly true for Quantinuum’s experiments, which had high enough gate fidelity to maintain a Linear Cross-Entropy score of order 1 (i.e., they’re no longer all that “noisy”). The central drawback of these experiments is no longer lack of confidence about quantum advantage; rather, it’s just that we only get samples as output, and directly verifying the quality of the samples seems just as intractable for a classical computer as spoofing the samples.
  3. As of this past year, however, we have some strong candidates for verifiable quantum advantage. One is the Google OTOC experiment, as even Hagar himself acknowledges (that’s his “single clear exception”). A second is the simulations of the 2D Fermi-Hubbard model on Quantinuum and Google machines, like this one. The 1D Fermi-Hubbard model can be classically simulated pretty easily (see here for example), but the 2D one still presents challenges, meaning that in some regimes, the best available estimates of certain observables apparently now come from quantum computers. I wish I could write about other examples that will be public shortly.
  4. Yes, the “real” goal remains, as it’s been since the 1990s, to build a scalable fault-tolerant quantum computer—and I’m glad that Hagar (unlike, say, Gil Kalai) never suggests that we’ve learned anything to rule that goal out. In the meantime, an intermediate goal would be to use NISQ devices to do physics and chemistry simulations that are commercially useful, or that help solve important scientific problems. The point of quantum supremacy experiments, you might say, is that by demonstrating the reality of quantum speedup about as clearly as it can be demonstrated with current hardware, they let us cleanly turn our attention to those more ambitious goals.

Anyway, my son and I need to catch a plane to Utah now, for the next iteration of the wonderful Epsilon Camp, where I’ll again be teaching theoretical computer science to 11- and 12-year-olds. But feel free to discuss in the comments! Nothing about world affairs in this thread please, just quantum supremacy.

Update (July 19): Not unrelated to the subject of this post, here’s a podcast I did with Gill Eapen of “Scientific Sense” about the current situation in quantum computing including recent experimental victories.

July 18, 2026

John PreskillMy friend Mark Wise

Mark Wise, the John A. McCone Professor of High Energy Physics at Caltech, passed away on July 10 at age 72. At a recent memorial service, John Preskill made these remarks.

I’m John Preskill, Mark’s colleague on the Caltech physics faculty for more than four decades. Our friendship goes back even farther. My wife Roberta and I met Mark and Jackie not long after they arrived at Harvard in 1980. We’ve been friends since then. We attended the bris for both Barry and Jonathan during those Harvard days. Mark and Jackie have two boys and we have two girls who are a few years younger, who were thrilled to connect with Barry and Jonathan when the families would get together for occasions like Passover or Hanukkah or Thanksgiving. When the kids were little, Mark and I would sometimes muse about the potential for forging even closer family ties if those relationships blossomed.

That didn’t happen. But Mark would preside at each Seder with a light hand, sprinkling the occasion with corny jokes as was his style, and Jackie would be determined to make it to the end of that customized family-friendly Haggadah she had meticulously prepared. The children, meanwhile, would be wondering when they’d be able to continue their game of sock baseball.

Many of you know that Mark was deeply dedicated to his family and friends. I’ll make some brief remarks about three facets of Mark I know especially well: Mark the scientist, Mark the teacher and mentor, and Mark the colleague and friend.

Because of his self-deprecating manner, those of you who are not scientists may not appreciate Mark’s stature as a physicist. He was one of the most influential figures in theoretical particle physics of his generation. It was not obvious things would turn out that way. Growing up in Toronto, Mark was an indifferent student, and his poor grades reflected that lack of interest. As a 9th grader, though, it struck him that he better change his ways and figure out how to make something of his life. He liked sports — the possibility of being a professional athlete was briefly considered, but discarded. Somehow he decided that science would be a better fit. I’m not sure why — he had recently failed math. But he worked hard and had inspiring teachers, so by the time he finished high school Mark was an excellent student, and he sailed into the University of Toronto well prepared to major in physics,

At U of T, Mark came under the influence of a young professor, Nathan Isgur, who would later become his close research collaborator. Under Nathan’s guidance, Mark sought admission to the PhD programs of the most prestigious US research universities, intent on a career devoted to deep exploration of the fundamental laws of physics. He was rejected everywhere he applied. He should have been discouraged. But he wasn’t. Mark shrugged and said: “It’s okay. I’ll stay another year in Toronto, I’ll get a master’s degree, I’ll apply again and I’ll get in somewhere.” And that’s what happened. He went to Stanford, where, under the kind tutelage of Fred Gilman, Mark took off like a rocket. Hired to the Caltech faculty in 1982, he was a tenured full professor three years later at the age of 31, and appointed as the John A. McCone Professor of High Energy Physics while still in his 30s.

Mark liked action movies, such as those starring Arnold Schwarzenegger or Clint Eastwood. In serious moments, we would sometimes ponder together why we’re successful at what we do, and Mark would always quote Clint Eastwood as Harry Callahan in Magnum Force: “A man’s got to know his limitations.” We would both laugh, but those were words of wisdom. Mark understood what he did well as a research scientist and what he was less good at. Finding problems he could solve that would have interesting consequences for experiments that had been done or could be done was where he excelled – he did it again and again. Mark never lost his zest for calculating things, often by hand with pen and paper, his head resting on one arm with his glasses pushed up onto his forehead as he scribbled. Getting to an answer that was experimentally relevant never stopped giving him a thrill.

Mark also never lost his sense of appreciation for the teachers and mentors who had inspired and helped him. Perhaps that’s why he became such a dedicated teacher and mentor himself. It’s hard to impress Caltech students, but Mark’s lectures where extremely popular, not just for their pedagogical value but also for the humanity and humor he displayed. Students had to pay attention because otherwise one might miss the jokes, which inevitably became known as “Wisecracks.” There is even an account on X with the handle @MarkWiseSays, curated by students who want to preserve Mark’s pithy lessons in physics and in life.

For example, Mark might say: “If you really get depressed, I recommend diagonalizing a 2×2 matrix.” For physics students, this is both funny and sage advice. Or he might say. “This calculation will knock your socks off.” A cliché you might hear from anyone. But who besides Mark would then proceed to remove his shoes, rip off his socks, hurl them at the blackboard, put his shoes back on and resume lecturing?

Most famously, Mark would come to class with an ample supply of coins. He would ask the class questions, sometimes about physics and sometimes random trivia, rewarding a student who gave an answer Mark approved of by tossing a coin. At first the coins were quarters. But Mark, who had a scholarly interest in finance as well as physics, eventually felt that due to inflation he needed to upgrade to dollar coins. These are harder to come by, so it took frequent visits to the bank to make sure he wouldn’t run out. His antics in class made Mark human and approachable, and students responded. Mark felt that many Caltech students don’t fully realize how smart they are. He saw part of his job as building their self-confidence and relieving their stress.

As a colleague and mentor to graduate students and postdoctoral scholars, Mark was highly collaborative. He believed that interactions with others sparked his creativity. He was never at all pompous. I know this started early. When we were in the Harvard Society of Fellows we were obligated to have dinner with the Senior Fellows on Monday nights. It was a rather stuffy occasion. And, though I don’t think they do this anymore, after a sumptuous meal we would literally retire for brandy and cigars. Once, while puffing on his cigar after dinner, Mark had an inspiration. He gathered up a few junior fellows and led them to a theater for a movie he thought everyone should see right away. The movie was Conan the Barbarian. And everyone had a blast. That was a perfect Mark moment.

As the news about Mark has spread, accolades have poured in from physicists all over the world. He was admired not just for his scientific brilliance, but almost as much for his quirky sense of humor and his kindness. Mark was a wonderful friend to many of us. When you were with him, you were sure to laugh and feel good. He touched the lives of countless colleagues, students and friends. We miss him terribly but there are so many memories that we’ll cherish. We are all so very fortunate to have known and loved Professor Mark Wise.

Photo credit: Clara Murgui, Caltech

Doug NatelsonA few optics/metamaterials highlights from META 2026

This past week I attended META 2026, the 16th International Conference on Metamaterials, Photonic Crystals and Plasmonics, at Trinity College, Dublin.  This was the first time I've ever gone to this conference, which has grown from somewhat blurry beginnings to a ~ 900+ person annual event.  Here are a few scientific highlights:
  • Metasurfaces, built up from spatial arrays of dielectric (or sometimes semiconductor or plasmonic) resonators called "meta-atoms", have matured into very impressive, versatile tools.  In her plenary talk, Ruwen Peng from Nanjing showcased different approaches, combining angularly rotated meta-atoms ("Pencharatnam-Berry") and size-modulated meta-atoms.  The result can produce polarization-entangled photon beams, entangle photon spin and orbital angular momentum for quantum key distribution, and do full entanglement distribution over many channels.  Similarly, Federico Capasso gave a very impressive talk about the progress in the field, from visible wavelength flat optics ten years ago to compact platforms for sophisticated quantum tomography.
  • Nikolay Zheludev gave a great overview about combining measurements + machine learning estimators (e.g., here) to achieve effective optical resolution far better than conventional limits.  This can be used to make optics-base estimates of nanowire lateral displacements down to the 100 pm level, for example.  Rather than looking at the flow of energy in an optical imaging system, one can look at the flow of Fisher information regarding the object being imaged.
  • There were a series of talks throughout the meeting about chirality of optical scattering, what this means, and what it can lead to (including enantiomer-selective imaging and chemistry).  Note that it's important to distinguish between intrinsic chirality (e.g., the object scattering the light has a real structural handedness), extrinsic chirality (the object scattering the light is not chiral, but the experimental arrangement to do and measure the scattering introduces chirality into the measurement), and chirality in the fields themselves (think swirling Poynting vectors locally) that don't necessarily extend to the far field.  There are some neat probes of local effects, like this use of local polymerization.
  • Roman Quidant gave a talk about metalenses that are also optomechanical structures (e.g., use a pump beam to excite mechanical deformation of the metalens to steer the focus of a probe beam).  This lets you do some pretty neat things, like control the sign of optical forces by dynamically tuning the relative importance of momentum transfer (pushing objects with light by direct momentum kick from photons) and polarization forces (the classical optical tweezer situation where polarizable objects "seek" regions of high intensity).  This can enable feedback control to do optical cooling of trapped, levitated particles, potentially down to the quantum level.
  • Alessandra Boltasseva presented a variety of recent advances, including a look at how plasmonic ceramics like TiN and HfN have properties that can be dramatically tuned as their thickness gets down to the few-unit-cell level, a regime she and collaborators term "transdimensional" (to distinguish from atomically thin 2D van der Waals materials).  The possibility of Wigner crystallization in such systems is exciting, though disorder is a likely complication.  
  • A 4-channel wavelength division multiplexer
    made from etched Si3N4, from this paper.
    Jeremy Baumberg talked about building metamaterials out of molecularly-spaced nanoparticles, and how this has opened up real opportunities for chemical sensing based on surface-enhanced Raman and infrared absorption, as in this example.  Neat stuff.
  • There were multiple talks about metasurfaces for nonlinear optics, including one by Igal Brener on cool ways to use GaAs metasurfaces to produce entangled photon pairs via bound states in the continuum.  
  • Likewise, there were a number of presentations about inverse design, where computational tools are used to produce very funky looking structures which can act as, e.g., multichannel routers of optical signals.  Jelena Vuckovic presented an overview of this, showing how it can be done at scale to produce a chip that acts as a 1 TB/s optical router.  Structures produced this way always seem to me like some kind of eldritch geometry out of HP Lovecraft (see figure), but they work.  
As always, apologies to those whose work I didn't mention above; my note-taking was pretty uneven.

(I am trying to strike a balance between talking and educating people about science, which is basically the point of this blog, and keeping people informed/voicing some of my personal opinions about the crisis in the US research ecosystem (arguably the most consequential challenge facing US researchers today, with long-term implications that will be felt for many years).  There is still very exciting work being done in nanoscience, the physics of materials, etc. - we are just facing a future where if current trends continue the major advances may increasingly happen outside the US.)

July 17, 2026

Matt von HippelIn Defense of Reductionist Chauvinism

I don’t think people who argue about reductionism are really arguing about reductionism.

Reductionism is the idea that the behavior of big, complicated things (people, economies, ecosystems) boils down to the behavior of their smallest constituents (molecules, atoms, subatomic particles). It’s often contrasted with emergence, the idea that new rules emerge in those big, complicated systems that are more than just the rules that govern the smallest scales.

Emergence can be divided into two kinds: weak and strong. In strong emergence, the big, complicated things have their own causal powers that aren’t due to smaller things at all. This tends to get mystical, with ideas like “lifeforce” and “consciousness”. Weak emergence is much milder, and while the big complicated things are best described by their own laws, in weak emergence they are in principle still caused by laws on smaller scales.

In practice, basically no-one believes in strong emergence: it seems way too much like magic for most scientists. And basically everyone believes in weak emergence: it would be nuts to insist that economists and biologists aren’t discovering important rules that would be almost impossible to find for someone who just had physics and chemistry to work with.

So if everyone agrees, what do people argue about?

Arguments about reductionism are really arguments about attitudes. If you position yourself as a reductionist, or an emergentist, you’re defending a particular way of thinking about the world, one that privileges one scale over another. When people argue for emergence, what they really seem to be doing is opposing a kind of “reductionist chauvinism”, where people like physicists insist that their perspective is the most valuable one.

And that’s understandable, because physicists can definitely be jerks sometimes. Let it be known I am no fan of jerks.

But I think it’s worth defending reductionism, not as a philosophy, but as an attitude or perspective. Worth arguing not about whether things in practice reduce or not, but about whether reduction is a good goal, about whether science that reduces more successfully is healthier science.

Because I think it is. And the reason boils down to agreement.

Simpler systems are less controversial systems. When we write down the laws that govern subatomic particles, they’re more definite: less heuristic, more precisely specified, with fewer exceptions. That isn’t to say there’s zero controversy in these subjects, there’s even controversy between mathematicians. But the more thoroughly you can boil something down to simple rules, the easier time you have of convincing others you’re right.

In contrast, the laws of the largest scales, like psychology and biology, are deeply heuristic. They thrive on exceptions and guidelines, general tendencies without clearly defined limits. And the problem with such laws is that they can lead to intractable arguments. Different schools of thought in psychology may simply never be able to convince each other, and may just have to wait for one to die off, the aesthetic feel of one set of ideas falling out of fashion as people become preoccupied with different sorts of problems.

Every time you can reduce, you avoid those insoluble disagreements. The simpler a system you can invoke, the more you can cooperate and build off each other’s work, the less time you have to waste disagreeing, the more you can accept people with different aesthetic and philosophical preferences as just different perspectives on what are ultimately the same facts. Reductionism is a technology for peace, and one of the most powerful we have.

So yeah, I’m a bit of a chauvinist about reductionism. That’s typical of an ex-physicist, sure. But it comes from a place of concern. I want a peaceful world, where we can learn from each other and find ways to agree. And reductionism is how I get there.

Jordan EllenbergI saw the figure 7 in black

Dream. I’m giving some kind of lecture to grad students with not fully prepared slides. There are slides that I made some time ago and have forgotten whe I meant by them. But what feels important is an image that’s just the number “7”, in black, sans serif font, against a plain white background; this is meant to suggest an unequivocal decision, something that just is a certain way, without ambiguity.

When I woke up this morning I found this dream and its central image puzzling. But it’s actually very transparent. I had to decide when to have surgery to repair my torn rotator cuff. Oh did I forget to mention that when I knocked my shoulder out of the socket I also tore my rotator cuff? So it has transpired. Apparently this happens a lot.

I worried a lot over the optimal date to schedule the surgery. Eventually I decided on August 7. The other option was today.

Terence TaoTwo more apps: visualizing the zeta process and the motions of the heavens

I believe that the creation of visualization apps to illustrate mathematical or scientific concepts is a particularly favorable use case for modern coding agents, as many of the downside risks attached to other LLM use cases are limited:

  1. Not mission-critical. As such apps are not authorative sources of truth and only used for secondary purposes, a small positive error rate in the output can be acceptable.
  2. Stand-alone. As the applets are not destined to be incorporated into a larger codebase or literature, the technical debt incurred by delegating all the coding to an LLM agent is bounded.
  3. End product is deterministic (and sandboxed). As the applets run on a deterministic language (Javascript), are sandboxed against file or internet access, and do not make any LLM calls at run-time, security and privacy concerns are minimal, and the applet can be maintained without continued premium LLM access or resource-intensive compute.
  4. Not replacing primary skills. While deskilling is the tradeoff one accepts when relying on these tools to accelerate output, I am perfectly willing to forego the opportunity to keep my Javascript skills at a high level, as this is a tertiary skill for me at best in my chosen profession. (I continue to manually program in Lean and in Python to keep in practice with programming in general.)
  5. Not competing with humans. To my knowledge, there is no existing human effort that is being duplicated by these applets (the activity in this direction appears to have peaked two decades ago).

I would however caution against unrestricted LLM use when one or more of the above five favorable situations is not in effect.

With these points in mind, I have used such an agent to create two further apps. The first app illustrates the “zeta process” that was introduced in my recent paper with Alexeev, Barreto, Li, Lichtman, Price, Shah, and Tang, though it was first discovered by an AI. For each s > 1, the zeta distribution Z_s is a random natural number with distribution

\displaystyle {\bf P}( Z_s = k ) = \frac{1}{\zeta(s) k^s}.

It has long been known that this distribution has good number-theoretic properties: for instance, the number of times a given prime p divides Z_s has a geometric distribution of mean p^{-s}. However, the new observation is that these random variables Z_s can be chained together into a single stochastic process, which we call the “zeta process”, which is an infinite divisibility chain. I used an agent to create an app to visualize this process:

The underlying process is generated by several exponential random variables at each prime: in the above instantiation of the process, two such variables are visible at the prime p=2, and one variable at the primes p=3,5. At a given choice of s, Z_s is formed by collecting all the variables below this threshold (and for which all predecessors also lie below the threshold); in the above illustration, this amounts to one variable at each of the primes p=2,3,5, leading to Z_2 = 30 in this case. Additional visualizations in the app display the distribution of each Z_s, as well as the distribution of the hitting probability \nu_\Lambda, which among other things can be used to give a quick solution to Erdős problem #1196.

The second app is rather different in nature, and is a somewhat whimsical attempt to display the motion of the heavens, both at “human” scales of space and time, and at more “astronomical” scales (in which the motion of the planets in particular are more apparent). It is very loosely inspired by the game “Katamari Damacy“, in which one absorbs both terrestrial and celestial objects of many different scales. Here is how the app typically looks at a human scale:

And here is how it looks when one’s perspective leaves the Earth’s atmosphere:

(As I did not want to render an entire explorable world in this app, the observer in the app is only limited to changing his or her size, from a human to a creature of comparable size to the Earth itself; they cannot move horizontally on the planet.) At the largest scales of space and time, the classic orrery diagram appears:

After lengthy conversations with the agent, I was able to implement many astronomical phenomena, including phases of the Moon, the effect of Earth’s rotation against the fixed stars (though one can also stabilize one’s view against those stars to see the Earth’s rotation more directly), and so forth.

July 16, 2026

Tommaso DorigoDefining AI

Defining AI

Artificial Intelligence will probably be remembered as the least well-defined technological advancement of humankind.

Tommaso Dorigo

Jordan EllenbergWalking Red Flag

I have trouble sleeping — trouble staying asleep, not trouble falling asleep. I have found that the best way to get back to sleep at three or four in the morning is to listen to a podcast on a 10-minute fadeout timer. But the podcast has to land in a very specific zone: interesting enough that I don’t actively mind listening to it, but not so interesting that my brain maintains awakeness in order to catch what’s next. After a long process of trial and error I have found that millennial dating advice podcasts are perfect. Getting these little glimpses into the lives of the young and confused is never dull, but at the same time, anything you miss is going to happen again to some other hapless millennial single in a week or two, so it’s OK to drift off.

Anyway: Jared Freid, one of the hosts of my go-to millennial dating podcast, U Up, has a book out, called Walking Red Flag. It’s really good, even while being of no practical relevance to me! I like it, Dr. Mrs. Q likes it, AB likes the parts I’ve read to her. Presumably actual lonely millennials like it too. Freid is a stand-up comic, and as such has a deep feel for words and sounds. Pardon this coarse language, but I will never forget him describing his small sad NYC apartment with the sentence “I fuck next to the sink.” That is writing! That’s what you learn in poetry class! If it’s a good poetry class.

And not to be corny, but as the father of a son, I appreciate that Freid can be out there being a straight man who both wants to date women and also clearly actually likes women and is relentlessly normal about it. Bizarrely little of this in the male advice space!

Anyway, this is probably the only comic millennial dating advice book I am ever going to recommend on this blog, but it really is — there’s no other way to put this — a hoot.

July 15, 2026

Terence TaoVisualizing the Gilbreath expectation sequence

One byproduct of learning how to use coding agents to create visualization apps is that it now becomes straightforward to convert any figure in one’s papers that had already been generated by code (e.g., in Python) into a more interactive, animated applet.

I can illustrate this with Figure 1 from my recent paper on the Gilbreath conjecture with Chase and Hunter, reproduced below:

This plot displays both exact and numerically simulated values of a certain poorly understood sequence c_n relating to the Gilbreath conjecture, which I will call the “Gilbreath expectation sequence” here for lack of a better name. The definition of the sequence is as follows. Consider a “Gilbreath array” which is an inverted pyramid, where the top entries are n+1 independent exponential random variables of mean 1, and all the other entries are the absolute values of the differences of the two entries immediately above it. Thanks to the visualizer app, I can quickly give an example (with n+1=6):

The left diagonal entries are then random variables; the sequence c_0,\dots,c_n are defined to be the expectation of these values. (The process is stationary, so in fact any entry on the i^{th} row will have expectation c_i.)

If one starts with the first n normalized prime gaps (which have expectation about \log n/2, and are conjecturally distributed asymptotically according to a geometric distribution), then standard conjectures (e.g., the prime tuples conjecture) predict that the k^{th} row entries should decay like c_k\log n/2, at least for small k. So the Gilbreath conjecture appears to be tied to how fast the sequence c_k decays with k.

One can in principle work out each value of c_k as an explicit rational number by performing a certain complicated multivariate integral, but in the paper we only did this for k \leq 3 (the orange line in the above figure); for the remaining k we performed a Monte Carlo simulation with 10^6 Gilbreath arrays to obtain a numerical approximation (in blue), which (as per the law of large numbers) agreed well with the theoretical values. A later calculation of Michael Ross extended the theoretical values to k \leq 6, maintaining the good fit:

The asymptotic behavior of the sequence c_n remains mysterious. Clearly, it is not monotonic; in fact we cannot even prove it is bounded. The best we could do in our paper was establish an inequality which, roughly speaking, showed that c_n cannot decay faster than 1/n.

In a recent preprint of Ross, these numerics were extended, and a rough empirical prediction

\displaystyle c_n \approx C \lambda^{s_2(n)} / n

was proposed for some constants C>0 and \lambda>1 (empirically \lambda \approx 1.17), where s_2(n) is the number of 1’s in the binary expansion of n; in particular, it is the fluctuation in this quantity s_2 that is intended to explain much of the non-monotonic behavior of c_n. These are now all displayed in the following companion applet, which was a routine matter to generate in about an hour by the coding agent (which by this point has extensive experience with creating such apps, encoded via a “skill” markdown file that it maintains):

The appearance of the quantity s_2(n) may initially appear mysterious, but it is related to Lucas’s theorem, Kummer’s theorem, and the Sierpinski gasket. Consider for instance a Gilbreath array where all the entries are zero except for a single “spike”. Then the following Sierpinski pattern emerges:

Here is what an n=64 version of this picture looks like (with the spike positioned at the 32th entry):

The number of 1s in the k^{th} row is then 2^{s_2(k)} (if we index the rows starting from zero), which is at least of the same shape as the empirical prediction, albeit with different constants. (This sequence is also known as Gould’s sequence.)

Numerically, we seem to observe fragments of Sierpinski gaskets being generated before decaying (often due to “collisions” with other gaskets):

However, it is not clear to me at all what the asymptotic probabilistic model should be, even heuristically; it does not resemble any random shape model that I am familiar with. But perhaps there are readers more expert in probability theory or statistical physics who may be able to suggest such an asymptotic limit?

July 14, 2026

Terence TaoCall for long programs, workshops, and summer schools at IPAM

(I am writing here in my capacity as Director of Special Projects at IPAM.)

IPAM seeks program proposals from the mathematical, statistical, and scientific communities for long programs, workshops, and summer schools.  Most program proposals are reviewed at IPAM’s Science Advisory Board meeting, held in November each year.  Programs are selected on the basis of their scientific impact and contribution to IPAM’s goals. IPAM is committed to supporting a community where people of all backgrounds and points of view can engage, learn, and thrive.  If you would like to discuss your program ideas and prepare a proposal for IPAM’s consideration, you are encouraged to contact the IPAM Director. For more information visit: https://www.ipam.ucla.edu/propose-a-program/long-programs-2/

Doug NatelsonBad to worse at NSF? (July 2026 edition)

When I wrote this post last month, I really did not want it to be first in a series.

The inciting incident to write that post was a news article in Science reporting that there have been draconian ~ 30% cuts in present fiscal year budgets within the NSF.  Program officers were instructed to keep this confidential and not talk about it with PIs.  The rumor was that this funding would go to support the TIP directorate and its activities, particularly the "X-Labs".  

Now Dan Garisto has broken this story in Nature. For some, this is behind a paywall, so let me hit the highlights.

  • "NSF staff members — who asked to remain anonymous out of fear of retaliation — and an internal NSF ledger seen by Nature suggest that the NSF plans to claw back around US$500 million that has already been distributed to grant-making divisions. "
  • "Several NSF staff members told Nature that at least some of the funds will be funnelled to another project, a brainchild of the White House OSTP."
  • "Even though programme officers have less to spend than they had expected, those who spoke to Nature are unsure how many new grants will make it out the door. “They could make it through if the process is allowed to work without interference or interruption,” one staff member says. But most of “the process is now a black box and unpredictable”. "
  • "NSF staff members estimate that if the withdrawals are finalized, hundreds more proposals that have been recommended for funding across the engineering, computer and information science and engineering, and maths and physical science directorates would need to be sent back to programme officers for revision. Some proposals would be held until a later date when funds are available. Others would have their budgets reduced, and some would simply be declined."
  • "Staff members say that they are frustrated by the planned diversion of funds and the lack of communication about the agency’s spending. “We don’t know where the money’s going or what’s going on,” says one staff member. Programme officers are not allowed to pass on what they know to researchers. “We cannot communicate to the community at all. We’re forbidden.”"
Basically, this reporting confirms the unprecedented mid-fiscal-year cuts; confirms that there is zero transparency about this with the community and even within the agency; and reveals that the funds are apparently being diverted to some mysterious OSTP project.  

To get a sense of scale, consider a typical program within MPS, with an annual budget of around X dollars.  About 20% or more of those funds are "mortgaged", set aside for ongoing already-funded projects, leaving 0.8X dollars for new awards.  Now, however, there is a 30% cut in the middle of the fiscal year, so the resources now available for new awards are 0.7X-0.2X = 0.5X.  The number of new awards therefore has to be cut by 37.5% relative to last year.   

This is a huge deviation from what Congress appropriated, seemingly being steered at the direction of the White House.  Will Congress care, or will the majority just give this power to the executive branch?  Is anyone going to disclose what the mystery project is?  Is any major national media outlet going to report on this?

I will post more science soon, I promise, and I hope that I don't have more posts in this series.

Terence TaoA paper diagram visualizer

I am finding the newly revealed capability to code old applet ideas into reality to be very tempting to sink more time into, though I am certainly encountering the common “vibe coding” experience that the process can produce something that superficially resembles a finished product well before a satisfactory level of testing and review has been completed; indeed, it is the review process which is now the most time-consuming, to the point where I think any further advances in coding agent capability will have little impact on the new bottlenecks in the design process.

In any event, I spent a few hours working to realize a proposal I had made back in 2023 to automatically create diagrams to visually illustrate the logical flow of a given mathematical paper. At the time, Freddie Manners, extrapolating from the half-decent capability of the then-newly released ChatGPT 3.5 at this task, presciently predicted that “by the time a dedicated tool had been completed, the next general purpose engine would be better than it”.

With that in mind, I decided to focus not on the generation of the diagram – which now can be done at various levels of quality by any number of large language models – but on its presentation. The result is the following app, which can take a certain formatted JSON file of dependencies between theorem objects and produce an interactive graph which can be explored, edited and also exported (somewhat lossily) into other standard formats such as SVG, TikZ, quiver, or Mermaid. Here is a screenshot of a diagramming of the celebrated proof by Wang and Zahl of the three-dimensional Kakeya conjecture:

Using an LLM, I generated diagrams for eight papers for demonstration purposes, including for instance a diagram for Wiles’s proof of Fermat’s last theorem, or of Szemeredi’s proof of his famous theorem on arithmetic progressions (which sports a notoriously convoluted such diagram in the original paper), as well as a few papers of my own. If there are other requests to diagram particular papers, I can try to use an LLM to generate more examples; but my intention is for users of the app to create their own such diagrams, either by manually constructing them, or by directing their own AI tools to build the diagram in the required format (which is a JSON, with the precise specification given here).

I mentioned in the previous post that for these sorts of visualization apps, which work deterministically for a given set of inputs, the downside risk of LLM use to build the app is acceptably low. For this particular app, there is a complicating factor, which is that while the app does remain deterministic, the data I used to populate the app – namely, the above diagrams – are also LLM-generated. I have done spot checks comparing the diagrams against the source papers, and did not find any errors; however, they are not guaranteed to be 100% accurate, and should only be used as approximations to the logical structure of these papers rather than completely exact representations. (The latter might become deterministically extractable should the results of these papers become formalized in a proof assistant language, but this is not currently the case.) Still, I hope these sorts of diagrams can serve as a helpful initial guide when first trying to read and understand a complex paper.

July 13, 2026

Scott Aaronson Held Prize call for nominations (+ call for postdocs)

Here at the National Academy of Sciences, it seems that my first job is to serve on the selection committee for the prestigious Michael and Sheila Held Prize in combinatorial and discrete optimization and related areas. The committee chair, my former MIT colleague Madhu Sudan (now at Harvard), invited me to share the following message here on Shtetl-Optimized. (I’d add: put in the effort to nominate someone, and you can actually influence how things go!)

Dear Colleagues

I am writing to seek nominations for the 2027 Michael and Sheila Held Prize. The scope of the prize and nomination needs are described below. If you intend to nominate someone I would appreciate a heads up by email to madhu@cs.harvard.edu one month before the deadline (so email by Sept 8, 2026) to let me know your nomination is coming. (We may also reach out to you in response to coordinate multiple/overlapping nominations.)

The Held prize honors outstanding, innovative, creative, and influential research in the areas of combinatorial and discrete optimization, or related parts of computer science, such as the design and analysis of algorithms and complexity theory. This $100,000 prize is intended to recognize recent work (defined as published within the last eight years, i.e., on or after October 6, 2018).

All nominations must be submitted online by Monday, October 5, 2026 and include:

1. Nomination letter describing the candidate’s work and why he or she should be selected for the award. No more than three (3) pages.

2. Curriculum vitae. No more than two (2) pages.

3. Bibliography listing no more than twelve (12) of the nominee’s most significant publications.

4. Suggested citation. A 50-word summary stating why the nominee should be considered for this award.

5. Two letters of support. No more than one letter of support can be written by someone of the same primary work institution as the nominee.

The Held Prize is given to a person or a set of persons, as supported by a paper or a body of work. Unless otherwise stated, preference will be given to scientists who may be earlier in their careers or those whose work has not been recognized by other prizes or awards. Nomination restrictions can be found here. Joint nominations will only be considered when nominees have collaborated closely on the paper to be recognized by the award. If nominating multiple individuals for a paper with additional authors, please clearly explain the reason for nominating those chosen, as well as the reason for excluding other collaborators, if applicable. 

Please feel free to circulate this call further within your department

Best
Madhu Sudan, on behalf of The Michael and Sheila Held Prize Selection Committee


And while I have your attention, a second CS theory announcement: David Soloveichik, my wonderful friend and colleague in UT Austin’s Electrical and Computer Engineering Department, has funding for a postdoc for 1-2 years, to work on the thermodynamics of computation here at UT. This is a topic that I’ve been trying to learn more about as well, so I might get involved too! David writes, “the big picture is to think of thermodynamics (energy dissipation / entropy production) as CS complexity measures like time and space usage.” If you’re on the postdoc market and this sounds potentially up your alley, email David to learn more.

July 12, 2026

n-Category Café Octonions and the Standard Model (Part 13)

When Lee and Yang suggested that the laws of physics might not be invariant under spatial reflection — that there’s a fundamental difference between left and right — Pauli was skeptical. In a letter to Victor Weisskopf in January 1957, he wrote:

“Ich glaube aber nicht, daß der Herrgott ein schwacher Linkshänder ist.”

(I do not believe that the Lord is a weak left-hander.)

But just two days after Pauli wrote this letter, Chien-Shiung Wu’s experiment confirmed that Lee and Yang were correct. There’s an inherent asymmetry in nature.

We can trace this back to how the ‘left-handed’ fermions and antifermions live in a different representation of the Standard Model gauge group than the right-handed ones. And when we try to build grand unified theories that take this into account, we run into the fact that while we can fit the Standard Model gauge group into Spin(10)\text{Spin}(10) in various ways, not all these ways produce the required asymmetry. There’s a way where it fits into Spin(9)\text{Spin}(9), which is too symmetrical to work… and alas, this one has a nice octonionic description!

To keep things simple I’ll explain this by focusing, not on the whole Standard Model gauge group, but its subgroup SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3). Here is a theorem proved by Will Sawin in response to a question of mine on MathOverflow:

Theorem 10. There are exactly two conjugacy classes of subgroups of Spin(10)\text{Spin}(10) that are isomorphic to SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3). One of them has a representative that is a subgroup of Spin(9)Spin(10)\text{Spin}(9) \subset \text{Spin}(10), while the other does not.

I’ll describe representatives of these two subgroups; then I’ll say a bit about how they show up in physics, and then I’ll show you Sawin’s proof.

We can get both subgroups in a unified way! There’s always an inclusion

SO(m)×SO(n)SO(m+n) \text{SO}(m) \times \text{SO}(n) \to \text{SO}(m+n)

and taking double covers of each group we get a 2-1 homomorphism

Spin(m)×Spin(n)Spin(m+n) \text{Spin}(m) \times \text{Spin}(n) \to \text{Spin}(m+n)

In particular we have

Spin(4)×Spin(6)Spin(10) \text{Spin}(4) \times \text{Spin}(6) \to \text{Spin}(10)

so composing with the exceptional isomorphisms:

Spin(4)SU(2)×SU(2),Spin(6)SU(4) \text{Spin}(4) \cong \text{SU}(2) \times \text{SU}(2), \qquad \text{Spin}(6) \cong \text{SU}(4)

we get a 2-1 homomorphism

k:SU(2)×SU(2)×SU(4)Spin(10) k \colon \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) \to \text{Spin}(10)

Now, there are three obvious ways to include SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) in SU(2)×SU(2)×SU(4)\text{SU}(2) \times \text{SU}(2) \times \text{SU}(4). There is an obvious inclusion

j:SU(3)SU(4) j \colon \text{SU}(3) \hookrightarrow \text{SU}(4)

but there are three obvious inclusions

,r,δ:SU(2)SU(2)×SU(2) \ell, r, \delta \colon \text{SU}(2) \hookrightarrow \text{SU}(2) \times \text{SU}(2)

namely the left one:

:SU(2) SU(2)×SU(2) g (g,1) \begin{array}{ccc} \ell \colon \text{SU}(2) &\to& \text{SU}(2) \times \text{SU}(2) \\ g & \mapsto & (g,1) \end{array}

the right one:

r:SU(2) SU(2)×SU(2) g (1,g) \begin{array}{ccc} r \colon \text{SU}(2) &\to& \text{SU}(2) \times \text{SU}(2) \\ g & \mapsto & (1,g) \end{array}

and the diagonal one:

δ:SU(2) SU(2)×SU(2) g (g,g) \begin{array}{ccc} \delta \colon \text{SU}(2) &\to& \text{SU}(2) \times \text{SU}(2) \\ g & \mapsto & (g,g) \end{array}

Combining these with our earlier maps, we actually get a one-to-one map from SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) to Spin(10)\text{Spin}(10). So we get three subgroups of Spin(10)\text{Spin}(10), all isomorphic to SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3):

  • There’s the left subgroup G G_\ell, which is the image of this composite homomorphism:

SU(2)×SU(3)×jSU(2)×SU(2)×SU(4)Spin(4)×Spin(6)kSpin(10) \text{SU}(2) \times \text{SU}(3) \stackrel{\ell \times j}{\longrightarrow} \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) \cong \text{Spin}(4) \times \text{Spin}(6) \stackrel{k}{\longrightarrow} \text{Spin}(10)

  • There’s the diagonal subgroup G δG_\delta, which is the image of this:

SU(2)×SU(3)δ×jSU(2)×SU(2)×SU(4)Spin(4)×Spin(6)kSpin(10) \text{SU}(2) \times \text{SU}(3) \stackrel{\delta \times j}{\longrightarrow} \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) \cong \text{Spin}(4) \times \text{Spin}(6) \stackrel{k}{\longrightarrow} \text{Spin}(10)

  • And there’s the right subgroup G rG_r, which is the image of this:

SU(2)×SU(3)r×jSU(2)×SU(2)×SU(4)Spin(4)×Spin(6)kSpin(10) \text{SU}(2) \times \text{SU}(3) \stackrel{r \times j}{\longrightarrow} \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) \cong \text{Spin}(4) \times \text{Spin}(6) \stackrel{k}{\longrightarrow} \text{Spin}(10)

The left and right subgroups are actually conjugate, but the diagonal one is truly different! We’ll prove this by taking a certain representation of Spin(10)\text{Spin}(10), called the Weyl spinor representation, and restricting it to those two subgroups. We’ll get inequivalent representations of SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3). This proves the two subgroups aren’t conjugate.

This argument is also interesting for physics. When restrict to the left subgroup, we get a representation of SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) that matches what we actually see for one generation of fermions! This is the basis of the so-called SO(10)\text{SO}(10) grand unified theory, which should really be called the Spin(10)\text{Spin}(10) grand unified theory.

(In fact this works not only for SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) but for the whole Standard Model gauge group, which is larger. I’m focusing on SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) just because it makes the story simpler.)

When we restrict the Weyl spinor representation to the diagonal subgroup, we get a representation of SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) that is not physically correct. Unfortunately, it’s the diagonal subgroup that shows up in several papers connecting the Standard Model gauge group to the octonions. I plan to say a lot more about this later.

The left subgroup

Let’s look at the left subgroup G G_\ell, the image of this composite:

SU(2)×SU(3)×jSU(2)×SU(2)×SU(4)Spin(4)×Spin(6)kSpin(10) \text{SU}(2) \times \text{SU}(3) \stackrel{\ell \times j}{\longrightarrow} \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) \cong \text{Spin}(4) \times \text{Spin}(6) \stackrel{k}{\longrightarrow} \text{Spin}(10)

Spin(10)\text{Spin}(10) has a 32-dimensional unitary representation called the ‘Dirac spinor’ representation. This representation is really on the exterior algebra Λ 5\Lambda \mathbb{C}^5. It’s the direct sum of two irreducible parts, the even grades and the odd grades:

Λ 5Λ even 5Λ odd 5 \Lambda \mathbb{C}^5 \cong \Lambda^{\text{even}} \mathbb{C}^5 \oplus \Lambda^{\text{odd}} \mathbb{C}^5

Physicists call these two irreducible representations ‘right- and left-handed Weyl spinors’, and denote them as 16\mathbf{16} and 16*\mathbf{16}\ast since they’re 16-dimensional and one is the dual of the other.

Let’s restrict the 16\mathbf{16} to the left subgroup G G_\ell and see what we get.

To do this, first we can restrict the 16\mathbf{16} along kk and get

214124* \mathbf{2} \otimes \mathbf{1} \otimes \mathbf{4} \; \oplus \; \mathbf{1} \otimes \mathbf{2} \otimes \mathbf{4}\ast

Here 1\mathbf{1} is the trivial representation of SU(2)\text{SU}(2), 2\mathbf{2} is the tautologous representation of SU(2)\text{SU}(2), and 4\mathbf{4} is the tautologous rep of SU(4)\text{SU}(4).

Then let’s finish the job by restricting this representation along ×j\ell \times j. Restricting the 4\mathbf{4} of SU(4)\text{SU}(4) to SU(3)\text{SU}(3) gives 31\mathbf{3} \oplus \mathbf{1}: the sum of the tautologous representation of SU(3)\text{SU}(3) and the trivial representation. Restricting 21\mathbf{2} \otimes \mathbf{1} to the left copy of SU(2)\text{SU}(2) gives the tautologous representation 2\mathbf{2}, while restricting 12\mathbf{1} \otimes \mathbf{2} to this left copy gives 11\mathbf{1} \oplus \mathbf{1}: the sum of two copies of the trivial representation. All in all, we get this representation of SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3):

2(31)(11)(3*1) \mathbf{2} \otimes (\mathbf{3} \oplus \mathbf{1}) \; \oplus \; (\mathbf{1} \oplus \mathbf{1}) \otimes (\mathbf{3}\ast \oplus \mathbf{1})

This is what we actually see for one generation of left-handed fermions and antifermions in the Standard Model! The representation 31\mathbf{3} \oplus \mathbf{1} describes how the left-handed fermions in one generation transform under SU(3)\text{SU}(3): 3 colors of quark and one ‘white’ lepton. The representation 3*1\mathbf{3}\ast \oplus \mathbf{1} does the same for the left-handed antifermions. The left-handed fermions form an isospin doublet, giving us the 2\mathbf{2}, while the left-handed antifermions have no isospin, giving us the 11\mathbf{1} \oplus \mathbf{1}.

This strange lopsidedness is a fundamental feature of the Standard Model.

The right subgroup would work the same way, up to switching the words ‘left-handed’ and ‘right-handed’. And by Theorem 10, the left and right subgroups must be conjugate in Spin(10)\text{Spin}(10), because now we’ll see one that’s not conjugate to either of these.

The diagonal subgroup

Consider the diagonal subgroup G δG_\delta, the image of this composite:

SU(2)×SU(3)δ×jSU(2)×SU(2)×SU(4)Spin(4)×Spin(6)kSpin(10) \text{SU}(2) \times \text{SU}(3) \stackrel{\delta \times j}{\longrightarrow} \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) \cong \text{Spin}(4) \times \text{Spin}(6) \stackrel{k}{\longrightarrow} \text{Spin}(10)

Let’s restrict the 16\mathbf{16} to G δG_\delta.

To do this, first let’s restrict the 16\mathbf{16} along k:SU(2)×SU(2)×SU(4)Spin(10)k \colon \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) \to \text{Spin}(10) and get

214124* \mathbf{2} \otimes \mathbf{1} \otimes \mathbf{4} \; \oplus \; \mathbf{1} \otimes \mathbf{2} \otimes \mathbf{4}\ast

as before. Then let’s restrict this representation along δ×j\delta \times j. The SU(3)\SU(3) part works as before, but what happens when we restrict 21\mathbf{2} \otimes \mathbf{1} or 12\mathbf{1} \otimes \mathbf{2} along the diagonal map δ:SU(2)SU(2)×SU(2)\delta \colon \text{SU}(2) \to \text{SU}(2) \times \text{SU}(2)? We get 2\mathbf{2}. So, this is the representation of G δG_\delta that we get:

2(31)2(3*1) \mathbf{2} \otimes (\mathbf{3} \oplus \mathbf{1}) \; \oplus \; \mathbf{2} \otimes (\mathbf{3}\ast \oplus \mathbf{1})

This is not good for the Standard Model. It describes a more symmetrical universe than ours, where both left-handed fermions and antifermions transform as doublets under SU(2)\text{SU}(2).

The fact that we got a different answer this time proves that G G_\ell and G δG_\delta are not conjugate in Spin(10)\text{Spin}(10). So to complete the proof of Theorem 10, we only need to prove

  1. Every subgroup of Spin(10)\text{Spin}(10) isomorphic to SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) is conjugate to G G_\ell or G δG_\delta.

  2. G δG_\delta is conjugate to a subgroup of Spin(9)Spin(10)\text{Spin}(9) \subset \text{Spin}(10), but G G_\ell is not.

I’ll prove 2, and then I’ll turn you over to Will Sawin to do the rest.

Why the diagonal subgroup fits in Spin(9)\text{Spin}(9)

Every rotation of n\mathbb{R}^n extends to a rotation of n+1\mathbb{R}^{n+1} that leaves the last coordinate fixed, so we get an inclusion SO(n)SO(n+1)\text{SO}(n) \hookrightarrow \text{SO}(n+1), which lifts to an inclusion of the double covers, Spin(n)Spin(n+1)\text{Spin}(n) \hookrightarrow \text{Spin}(n+1). Since we have exceptional isomorphisms

Spin(3)SU(2),Spin(4)SU(2)×SU(2) \text{Spin}(3) \cong \text{SU}(2), \qquad \text{Spin}(4) \cong \text{SU}(2) \times \text{SU}(2)

it’s natural to ask how the inclusion Spin(3)Spin(4)\text{Spin}(3) \hookrightarrow \text{Spin}(4) looks in these terms. And the answer is: it’s the diagonal map! In other words, we have a commutative diagram

SU(2) Spin(3) δ SU(2)×SU(2) Spin(4) \begin{array}{ccc} \text{SU}(2) & \xrightarrow{\sim} & \text{Spin}(3) \\ \delta \downarrow & & \downarrow \\ \text{SU}(2) \times \text{SU}(2) & \xrightarrow{\sim} & \text{Spin}(4) \end{array}

Now, we can easily fit this into a larger commutative diagram involving some natural maps Spin(m)×Spin(n)Spin(m+n)\text{Spin}(m) \times \text{Spin}(n) \to \text{Spin}(m+n) and Spin(n)Spin(n+1)\text{Spin}(n) \to \text{Spin}(n+1):

SU(2) Spin(3) Spin(3)×Spin(6) Spin(9) δ SU(2)×SU(2) Spin(4) Spin(4)×Spin(6) Spin(10) \begin{array}{ccccccc} \text{SU}(2) & \xrightarrow{\sim} & \text{Spin}(3) & \to & \text{Spin}(3) \times \text{Spin}(6) & \to & \text{Spin}(9) \\ \delta \downarrow & & \downarrow & & \downarrow & & \downarrow \\ \text{SU}(2) \times \text{SU}(2) & \xrightarrow{\sim} & \text{Spin}(4) & \to & \text{Spin}(4) \times \text{Spin}(6) & \to & \text{Spin}(10) \end{array}

We can simplify this diagram using the isomorphism Spin(6)SU(4)\text{Spin}(6) \cong \text{SU}(4):

SU(2)×SU(4) Spin(9) δ×1 SU(2)×SU(2)×SU(4) Spin(10) \begin{array}{ccccccc} \text{SU}(2) \times \text{SU}(4) & \to & \text{Spin}(9) \\ \delta \times 1 \downarrow & & \downarrow \\ \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) & \to & \text{Spin}(10) \end{array}

and then we can use our friend the inclusion j:SU(3)SU(4)j \colon \text{SU}(3) \to \text{SU}(4):

SU(2)×SU(3) 1×j SU(2)×SU(4) Spin(9) δ×1 SU(2)×SU(2)×SU(4) Spin(10) \begin{array}{ccccccc} \text{SU}(2) \times \text{SU}(3) & \xrightarrow{1 \times j} & \text{SU}(2) \times \text{SU}(4) & \to & \text{Spin}(9) \\ & & \delta \times 1 \downarrow & & \downarrow \\ & & \text{SU}(2) \times \text{SU}(2) \times \text{SU}(4) & \to & \text{Spin}(10) \end{array}

This shows that the diagonal subgroup G δG_\delta of Spin(10)\text{Spin}(10) is actually a subgroup of Spin(9)\text{Spin}(9)!

Why the left subgroup does not fit in Spin(9)\text{Spin}(9)

The three-fold way is a coarse classification of irreducible complex representations of compact Lie group. Every such representation is of one and only one of these three kinds:

1) not self-dual: not isomorphic to its dual,

2a) orthogonal: isomorphic to its dual via an invariant nondegenerate symmetric bilinear form, also called an orthogonal structure,

2b) symplectic: isomorphic to its dual via an invariant nondegenerate antisymmetric bilinear form, also called a symplectic structure.

I’ve written about how these three cases are related to the division algebras ,\mathbb{C}, \mathbb{R} and \mathbb{H}, respectively:

A complex representation is orthogonal iff it’s the complexification of a representation on a real vector space, and symplectic iff it’s the underlying complex representation of a representation on a quaternionic vector space.

But we don’t need most of this yet. For now we just need to know one fact: when nn is odd, every irreducible representation of Spin(n)\text{Spin}(n), and thus every representation of this Lie group, is self-dual: that is, isomorphic to its dual. In particular this is true of Spin(9)\text{Spin}(9).

Why does this matter? Assume the left subgroup G Spin(10)G_\ell \subset \text{Spin}(10) is a subgroup of Spin(9)\text{Spin}(9). When we restrict the Weyl spinor representation of Spin(10)\text{Spin}(10) to Spin(9)\text{Spin}(9) it will be self-dual, like every representation of Spin(9)\text{Spin}(9). Then when we restrict this representation further to SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) it must still be self-dual, since the restriction of a self-dual representation is clearly self-dual.

However, we know this representation is

2(31)(11)(3*1) \mathbf{2} \otimes (\mathbf{3} \oplus \mathbf{1}) \; \oplus \; (\mathbf{1} \oplus \mathbf{1}) \otimes (\mathbf{3}\ast \oplus \mathbf{1})

and this is not self-dual, since 1*1\mathbf{1}\ast \cong \mathbf{1} and 2*2\mathbf{2}\ast \cong \mathbf{2} but 3*3\mathbf{3}\ast \ncong \mathbf{3}.

So, it must be that G G_\ell is not a subgroup of Spin(9)\text{Spin}(9).

Proof of Theorem 10

To complete the proof of Theorem 10 we just need to see why there are just two conjugacy classes of subgroups of Spin(10)\text{Spin}(10) isomorphic to SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3). But in fact Will Sawin proved a stronger result! He was answering this question of mine:

Define the Standard Model gauge group to be S(U(2)×U(3))\text{S}(\text{U}(2) \times \text{U}(3)), the subgroup of SU(5)\text{SU}(5) consisting of block diagonal matrices with a 2×22 \times 2 block and then a 3×33 \times 3 block. (This is isomorphic to the quotient of U(1)×SU(2)×SU(3)\text{U}(1) \times \text{SU}(2) \times \text{SU}(3) by the subgroup of elements (α,α 3,α 2(\alpha, \alpha^{-3}, \alpha^2) where α\alpha is a 6th root of unity.)

Up to conjugacy, how many subgroups isomorphic to the Standard Model gauge group does Spin(10)\text{Spin}(10) have?

This question is relevant to grand unified theories of particle physics, as explained here:

This paper focuses on one particular copy of S(U(2)×U(3))\text{S}(\text{U}(2) \times \text{U}(3)) in Spin(10)\text{Spin}(10), given as follows. By definition we have an inclusion S(U(2)×U(3))SU(5)\text{S}(\text{U}(2) \times \text{U}(3)) \hookrightarrow \text{SU}(5), and we also have an inclusion SU(5)Spin(10)\text{SU}(5) \hookrightarrow \text{Spin}(10) because for any nn we have an inclusion SU(n)SO(2n)\text{SU}(n) \hookrightarrow \text{SO}(2n), and SU(n)\text{SU}(n) is simply connected so this gives a homomorphism SU(n)Spin(2n)\text{SU}(n) \hookrightarrow \text{Spin}(2n).

However I think there is also an inclusion S(U(2)×U(3))Spin(9)\text{S}(\text{U}(2) \times \text{U}(3)) \hookrightarrow \text{Spin}(9), studied by Krasnov:

Composing this with Spin(9)Spin(10)\text{Spin}(9) \hookrightarrow \text{Spin}(10), this should give another inclusion S(U(2)×U(3))Spin(10)\text{S}(\text{U}(2) \times \text{U}(3)) \hookrightarrow \text{Spin}(10), and I believe this one is ‘truly different from’ — i.e., not conjugate to — the first one I mentioned.

So I believe my current answer to my question is “at least two”. But that’s not good enough.

Sawin’s answer relies heavily on the 3-fold way — that’s why I told you that stuff about orthogonal and symplectic representations. When we embed the group SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) in Spin(10)\text{Spin}(10), we are automatically giving this group an orthogonal 10-dimensional representation, thanks to the map Spin(10)SO(10)\text{Spin}(10) \to \text{SO}(10). We can classify the possibilities.

He writes:

There are infinitely many embeddings. However, all but one of them is “essentially the same as” the one you studied as they become equal to the one you studied on restriction to SU(2)×SU(3)\text{SU}(2)\times \text{SU}(3). The remaining one is the one studied by Krasnov.

I follow the strategy suggested by Kenta Suzuki.

SU(3)\text{SU}(3) has irreducible representations of dimensions 1,3,3,6,8,6,10,101,3,3,6,8,6, 10, 10, and higher dimensions. The 1010-dimensional ones are dual to each other, as are the 66-dimensional ones, so they can’t appear. The 33-dimensional ones are dual to each other and can only appear together. So the only 1010-dimensional self-dual representations of SU(3)\text{SU}(3) decompose as irreducibles as 8+1+18+1+1, 3+3+1+1+1+13+3+1+1+1+1, or ten 11s. All of these are orthogonal because the 8-dimensional representation is orthogonal. However, the ten 11s cannot appear because then SU(3)\text{SU}(3) would act trivially.

A representation of SU(3)×SU(2)\text{SU}(3) \times \text{SU}(2) is a sum of tensor products of irreducible representations of SU(3)\text{SU}(3) and irreducible representations of SU(2)\text{SU}(2). Restricted to SU(3)\text{SU}(3), each tensor product splits into a sum of copies of the same irreducible representation. So SU(2)\text{SU}(2) can only act nontrivially when the same representation appears multiple times. Since the 3+33+3 is two different 33-dimensional representation, only the 11-dimensional representation can occur twice. Thus, our 10-dimensional orthogonal representation of SU(3)×SU(2)\text{SU}(3) \times \text{SU}(2) necessarily splits as either the 88-dimensional adjoint repsentation of SU(3)\text{SU}(3) plus a 22-dimensional orthogonal representation of SU(2)\text{SU}(2) or the 66-dimensional sum of standard and conjugate [i.e., dual] representations of SU(3)\text{SU}(3) plus a 44-dimensional orthogonal representation of SU(2)\text{SU}(2). However, SU(2)\text{SU}(2) has a unique nontrivial representation of dimension 22 and it isn’t orthgonal, so only the second case can appear. SU(2)\text{SU}(2) has representations of dimension 1,2,3,41,2,3,4 of which the 22 and 44-dimensional ones are symplectic and so must appear with even multiplicity in any orthogonal representation, so the only nontrivial 44-dimensional orthogonal ones are 2+22+2 or 3+13+1.

So there are two ten-dimensional orthogonal representations of SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) that are nontrivial on both factors, those being the sum of two different 33-dimensional irreducible representations of SU(3)\text{SU}(3) with either two copies of the two-dimensional irreducible representation of SU(2)\text{SU}(2) or the three-dimensional and the one-dimensional irreducible representation of SU(2)\text{SU}(2). The orthogonal structure is unique up to isomorphisms, so these give two conjugacy classes of homomorphisms SU(2)×SU(3)SO(10)\text{SU}(2) \times \text{SU}(3) \to SO(10) and thus two conjugacy classes of homomorphisms SU(2)×SU(3)Spin(10)\text{SU}(2) \times \text{SU}(3) \to \text{Spin}(10). The first one corrresponds to the embedding you studied while only the second one restricts to Spin(9)\text{Spin}(9) so indeed these are different.

To understand how to extend these to S(U(2)×U(3))\text{S}(\text{U}(2) \times \text{U}(3)), I consider the centralizer of the representation within Spin(10)\text{Spin}(10). Since the group is connected, this is the same as the centralizer of its Lie algebra, which is therefore the inverse image of the centralizer in SO(10)\text{SO}(10). Now there is a distinction between the two examples because the example with irrep dimensions 3+3+2+23+3+2+2 has centralizer with identity component U(1)×SU(2)\text{U}(1) \times \text{SU}(2) while the example with irrep dimensions 3+3+3+13+3+3+1 has centralizer with identity component U(1)\text{U}(1). In the second case, the image of U(2)×U(3)\text{U}(2) \times \text{U}(3) must be the image of SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) times the centralizer of the image of SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3), so this gives a unique example, which must be the one considered by Krasnov.

In the first case, we can restrict attention to a torus U(1)×U(1)\text{U}(1) \times \text{U}(1) in SU(2)×SU(2)\text{SU}(2) \times \text{SU}(2). The center of S(U(2)×U(3))\text{S}(\text{U}(2) \times \text{U}(3)) maps to a one-dimensional subgroup of this torus, which can be described by a pair of integers. Explicitly, given a two-by-two-unitary matrix AA and a three-by-three unitary matrix BB with det(A)det(B)=1\det(A) \det(B) =1, we can map to U(5)\text{U}(5) by sending (A,B)(A,B) to Aγ aBγ bA \gamma^a \oplus B \gamma^b where γ=det(A)=det(B) 1\gamma = \det (A) = \det(B)^{-1}, and then map from U(5)\text{U}(5) to SO(10)SO(10). This lifts to the spin group if and only if the determinant in U(5)\text{U}(5) is a perfect square. The determinant is γ 1+2a1+3b=γ 2a+3b\gamma^{ 1 + 2a - 1 + 3b} = \gamma^{2a+3b} so a lift exists if and only if bb is even.

The only possible kernel of this embedding is the scalars. The scalar A=λ 3I 2,B=λ 2I 3A = \lambda^3 I_2, B = \lambda^{-2} I_3 maps to λ 3+6aI 2λ 2+6bI 3\lambda^{3+ 6a} I_2 \oplus \lambda^{-2 + 6b} I_3 and so the kernel is trivial if and only if gcd(3+6a,2+6b)=1\gcd(3+6a,-2 + 6b)=1.

However, there are infinitely many integer solutions a,ba,b to gcd(3+6a,2a+6b)=1\gcd(3+6a,-2a+6b)=1 with bb even (in fact, a random aa and even bb works with probability 9/π 29/\pi^2), so this gives infinitely many examples.


  • Part 1. How to define octonion multiplication using complex scalars and vectors, much as quaternion multiplication can be defined using real scalars and vectors. This description requires singling out a specific unit imaginary octonion, and it shows that octonion multiplication is invariant under SU(3)\mathrm{SU}(3).
  • Part 2. A more polished way to think about octonion multiplication in terms of complex scalars and vectors, and a similar-looking way to describe it using the cross product in 7 dimensions.
  • Part 3. How a lepton and a quark fit together into an octonion — at least if we only consider them as representations of SU(3)\mathrm{SU}(3), the gauge group of the strong force. Proof that the symmetries of the octonions fixing an imaginary octonion form precisely the group SU(3)\mathrm{SU}(3).
  • Part 4. Introducing the exceptional Jordan algebra 𝔥 3(𝕆)\mathfrak{h}_3(\mathbb{O}): the 3×33 \times 3 self-adjoint octonionic matrices. A result of Dubois-Violette and Todorov: the symmetries of the exceptional Jordan algebra preserving their splitting into complex scalar and vector parts and preserving a copy of the 2×22 \times 2 adjoint octonionic matrices form precisely the Standard Model gauge group.
  • Part 5. How to think of 2×22 \times 2 self-adjoint octonionic matrices as vectors in 10d Minkowski spacetime, and pairs of octonions as left- or right-handed spinors.
  • Part 6. The linear transformations of the exceptional Jordan algebra that preserve the determinant form the exceptional Lie group E 6\mathrm{E}_6. How to compute this determinant in terms of 10-dimensional spacetime geometry: that is, scalars, vectors and left-handed spinors in 10d Minkowski spacetime.
  • Part 7. How to describe the Lie group E 6\mathrm{E}_6 using 10-dimensional spacetime geometry. This group is built from the double cover of the Lorentz group, left-handed and right-handed spinors, and scalars in 10d Minkowski spacetime.
  • Part 8. A geometrical way to see how E 6\mathrm{E}_6 is connected to 10d spacetime, based on the octonionic projective plane.
  • Part 9. Duality in projective plane geometry, and how it lets us break the Lie group E 6\mathrm{E}_6 into the Lorentz group, left-handed and right-handed spinors, and scalars in 10d Minkowski spacetime.
  • Part 10. Jordan algebras, their symmetry groups, their invariant structures — and how they connect quantum mechanics, special relativity and projective geometry.
  • Part 11. Particle physics on the spacetime given by the exceptional Jordan algebra: a summary of work with Greg Egan and John Huerta.
  • Part 12. The bioctonionic projective plane and its connections to algebra, geometry and physics.
  • Part 13. Two ways to embed SU(2)×SU(3)\text{SU}(2) \times \text{SU}(3) in Spin(10)\text{Spin}(10), and their consequences for particle physics.

July 10, 2026

Scott Aaronson Announcing BQP Partners: my and my brother’s new angel-investing venture

As I’ve written before, these past couple years I’ve often felt like the last remaining person in either quantum computing or AI who lacked a stake in some startup company whose valuation is right now shooting into interstellar space. My academic colleagues, including the ones who seemed the most singleminded about quantum oracle separations and other gloriously useless pursuits? One by one, like in a zombie movie, I learn that they too have now launched startups, and invariably raised tens of millions of dollars, for the sorts of ideas we might’ve idly traded at coffee breaks back in the day, before getting back to our real work.

So why didn’t I join this rollicking party? Partly because of a lifelong fear that, the instant my self-worth became tied to how much money I made, I’d need to humble myself before people who bluster and bully and lie and hype and conceal … yet who nevertheless succeed at becoming orders of magnitude richer than me. I’ve been terrified of even starting down that road, of whether I’d still be myself at the end of it.

It’s also partly that I can’t stand failure, or regret, or being wrong. Of course, as an academic researcher I also fail, and regret things, and am wrong constantly—but there it feels tolerable, because normally I can tell myself that it’s all just down to my inborn limitations. After all, if I could’ve solved the major open problem that someone else solved, or written the brilliant book that someone else wrote, then presumably I would’ve done it!

Clearly, though, I could’ve mined bitcoin in 2010. I could’ve gotten an early stake in Amazon or Google. It’s not even like those ideas never crossed my mind. I just … didn’t act on them, for some reason. (But even if I had, I’d probably just be full of regret that I hadn’t done even more.) Thus, my only way to avoid paralyzing regrets, has been to tell myself constantly that I’m not in the forecasting or money-making businesseses in the first place.

It helped that, insofar as I’m shallow or covetous, insofar as I’ve desired things of this world rather than insight or eternal truth, it’s never really been money that I cared about, but just being respected and liked. Elon Musk is the richest man on earth, but also one of the most despised—which isn’t a bargain that I could imagine ever appealing to me.

Plus, when I actually meet billionaires, I don’t find myself envious of their mansions or cars or anything else that they have; I don’t feel like such things would make my life any happier. Maybe I slightly envy their ability to fund the causes they care about, or their professional staffs who relieve them of drudgery, but mostly I envy the way their wealth announces, to whatever extent it does: “I was right when others weren’t.” Again, though, I’ve never trusted the world to cause me to be right about the future valuations of companies or anything similar, so I’ve settled for having been right about PostBQP and algebrization and BosonSampling.

The bottom line is that I made a choice decades ago to forgo trying to get rich, no matter how many of my friends did the same, and to strive instead to discover and tell the truth—to be a professor, a blogger, a jokester, and an “objective” arbiter and commentator. “Then, surely, everyone will like me!” my internal monologue went. “Then, surely, they’ll be grateful for all the free service I’ve rendered them—for decades of blogging, without once so much as asking for a donation or running an ad!”


HAHAHAHAHAHA.

As any regular reader will know, my attempts to be loved as a blogger backfired pretty spectacularly. Or rather: they did lead to thousands of strangers liking me (and I’m grateful for every last one of you), but they also led to probably an order of magnitude more strangers hating me, and congregating on Reddit and Twitter and elsewhere to discuss how badly I suck. And of course, trying to shift that balance by writing what people want to hear, rather than what I actually believe, was never within my realistic option set.

In the startup context, it didn’t matter how carefully I avoided taking a direct stake for or against any of the companies I blogged about. People on Twitter simply assumed that I had a stake—for example, that I must’ve shorted D-Wave or IonQ, or invested in their competitors, or had equity in AI companies. For why else would anyone write what I wrote?

Amusingly, my attackers here typically did have precisely the conflicts-of-interest that they falsely accused me of having, but that was never at issue; only my imaginary conflicts-of-interest were. Even as the Scott-haters greedily filled their pockets (or tried to), I alone needed to keep turning my pockets out to prove that they were still empty.


So then, screw it! In partnership with my brother David Aaronson, who’s long done investing professionally, and on David’s guidance and encouragement, I’m hereby embarking on a new policy.

Namely: when I hear about a brand-new startup that sounds relevant to my interests—in quantum, AI, or anything else—and I like and trust the founders (ideally, because of their previous academic research work), David and I will often make a small seed investment if the founders are open to it. Or, of course, we might become advisors or get involved in some other way.

In fact, David and I are launching BQP Partners—the link goes to our AngelList, where you can read about how to invest with us if you’re interested. (See also whether you can spot any differences between David’s writing style and preoccupations and mine!)

So far, David and I are investing in:

I have little doubt that more potential investments will come our way very soon (some, probably, as a direct result of this post).

Crucially, I can handle my burden of regret—the “why didn’t I do this much earlier, if I was going to do it at all?” question—by telling myself that friends of mine were not founding companies left and right until very recently. I can also tell myself that I’m doing this less as a bet about the future (in which case … what if I’m wrong?), than simply as a way to support brilliant colleagues doing things that I genuinely admire.

When I blog about a company, I’ll always disclose if I have a financial position that presents a clear conflict of interest, so you can judge for yourself whether to listen to me. (Although, if that’s the sort of thing you’d demand, then you probably weren’t listening to me in the first place, were you?)

Having reflected on it a lot these past few months, I’m happy with my new policy and with my and David’s new venture, and I’m curious to see where it goes. I’m at peace with the possibility that we’ll lose our shirts, but I’m even at peace with a more disturbing possibility—that we’ll make millions and then people will scream at me online for being a sellout, a hack, and a shill. Those people, as I’ve learned, were going to scream at me anyway.

Matt von HippelAmplitudes 2026, Part II

This is a continuation of my conference coverage from last week. The same warnings apply: this is much more technical than my usual posts, readers beware!

Last week I covered the talks from Monday through Wednesday, so I’ll jump right in here with Thursday morning, where the first speaker was Francesco Riva, who after a brief ad for his new board game Tutti Quantum gave a review of positivity constraints, first-principles restrictions on quantum field theories based on their behavior at high energies. After covering some of the research program’s successes like arguments against Galileons and massive gravity, he talked about how new methods allow one to take into account the possibility of loops of massless particles, essentially by invoking a formal version of the idea that real experiments have finite size. He was followed by Grant Remmen, who talked about his work deriving string theory-like amplitudes from increasingly minimal assumptions. While one can always quibble with the assumptions they impose, I do find it encouraging that they are now managing to do this game with both gravity and gauge theory, and with five-particle amplitudes, not just four.

Paul Heslop then covered applications of bootstrap techniques to non-supersymmetric theories, where analytic superspace still finds a way to be useful. Tomasz Taylor covered progress in calculating Yang-Mills amplitudes in de Sitter space. Andrea Puhm and Nima Arkani-Hamed don’t have slides online yet: Puhm’s title suggests her talk was part of the celestial holography field, while Nima’s was likely similar to his talk at Lancefest, where he talked about calculations of amplitudes in the limit of a very large number of loops (represented in the field with a capital L) or very large numbers of particles (represented in the field with a lower-case n). Given the different context, I’m guessing he left out the self-effacing jokes where he was “little n” and Lance was “big L”, though I’m hoping he at least mentioned his students had been checking their results against what they called the “lanswer”.

The evening ended with a gong show, which for the non-initiated is a series of short student talks with a strict time limit (hence the gong). I’m not quite so intrepid as to read all of the slides for these, commenters who attended are welcome to highlight special examples.

Friday began with a talk by Agnese Bissi, who reviewed the connection between holographic correlators and amplitudes in AdS space. I liked her emphasis on this as a lab to find nice new representations of amplitudes, and her summary at the end of the current frontiers. Axel Kleinschmidt followed with a talk on one-loop string amplitudes, where he and his collaborators have gotten gradually more proficient at manipulating the rich structure of elliptic functions that make an appearance. Piotr Tourkine talked about his work using the S-matrix bootstrap to find scattering amplitudes in higher dimensions, a context where accounting for thresholds presented a new challenge. This is a problem people are approaching with genuine supercomputers, he quoted one calculation at 100,000 CPU hours. Lauren Williams is one of a small community of mathematicians who have been intrigued by the amplituhedron, her talk was a walk through a series of conjectures, some made by physicists, some by mathematicians, most with counterexamples found in the last few years.

Finally, Zvi Bern closed the conference with a talk on the frontier of amplitudes-based gravitational wave calculations, referring to it, probably to David Kosower’s annoyance, as 5PM. Some parts of this frontier have been calculated, but a few have proved hard going, bottlenecked by immensely challenging integrals, which go beyond the capabilities of publicly available codes. The new integrals have a variety of strange functions, including the elliptics and Calabi-Yaus I spent time on in my own career, as well as Heun integrals, which I imagine I will have to learn more about from the slides of Elliptics & beyond ’26, as I don’t remember people talking about them when I was in the field. New integration strategies have led to improvements in the public codes and seem to be making good progress, but Zvi highlighted that ideally they want to not need supercomputers at all, as they’ll need to go higher in loops to see effects from, for example, the deformability of neutron stars.

Amplitudes 2027 will be in Munich, with the summer school in Mainz under Stefan Weinzierl’s capable hands. I’m looking forward to seeing what the state of the art looks like then!

Peter Rohde Zinalrothorn (4,221m)

An album of GoPro headcam footage climbing Zinalrothorn (AD, 4,221m) in Switzerland.

Full album (49 videos): https://youtube.com/playlist?list=PLFMVEM4j3NZ0

July 09, 2026

Andrew JaffeWhat My Thirty-Year-Old Algorithm Taught an AI

As a scientist in the later stages of my career, the managerial and mentorship load has increased, leaving less time for math and programming. These technical activities were also why I wanted to be a scientist in the first place, and I regretted losing the time for this hands-on research. Formerly the core of my scientific work, these activities require sustained attention, hours at a time, a resource now in short supply.

Also like many scientists, I’ve watched the coming of artificial intelligence over the last few years with growing interest. The technical predecessors and underlying substructure of these large language models (LLMs) are neural networks, a computer technology that has already begun to revolutionise many scientific fields, including cosmology and astrophysics. I wrote about LLMs in my recent book, The Random Universe, but hadn’t really used them for my own work.

So I started using Anthropic’s Claude (no particular reason for this choice, except that some colleagues had had positive coding experiences with it), figured out how to wire it up at the command line, and started talking to it. (After, yes, paying a subscription fee.)

My first project with one of these newly-capable AIs was a minor reanalysis of our data from the Planck satellite, seeing how the cosmological inferences respond to small changes in the data, part of the work for a recent paper. I knew exactly what I needed to do, but it was a lot of plumbing: getting disparate bits of software written by other people to work together in a way different from their authors had intended. I figured I could do it in a few days of solid work.

Instead, I pointed Claude to the draft of the paper, along with publicly available Planck data and code repositories, and asked it to implement the paper’s algorithms with Planck’s data. A few hours later, there was code, alongside tables, figures, and lots of tests to make sure I could trust — and understand — the results. It wasn’t (we weren’t) just able to write and debug the code quickly, it was able to run it, again and again, making small tweaks to the inputs and the code itself, and to the figures it generated, now part of our recent paper.

Next was something more involved: colleagues and I have created a program called Almanac to analyse specific kinds of cosmological data. We wanted to apply Almanac to some new results, in a regime in which it hadn’t really been tested (a very small patch, around 1% of the total sphere of the sky). Almanac helps us measure a curve called the power spectrum, which I’ve written about before.

I pointed Claude to our code, our papers, and to the new data, and explained the problem. Even ensuring that the (poorly documented) data was in a form that our code could understand would have taken me a few hours, but Claude suggested and implemented a series of tests to ensure that everything was self-consistent.

Almanac is a Monte Carlo sampler: because we are trying to understand the probability distribution of matter in the Universe, using noisy and incomplete data, the answer to our questions can only be given as probability distributions. Almanac is essentially a very complicated random number generator, and you can do self-consistency checks to see whether it is producing random numbers with the right properties.

Almanac’s results failed these tests. Could we understand why? Could we fix it? Now, rather than just plumbing, I needed Claude to help me diagnose the problem. It took a while.

Was it a simple bug? We did a series of tests showing that Almanac does work, essentially perfectly, on simpler datasets covering much more of the sky. In fact, this gave me the opportunity to ask Claude to write some new software, based on a paper and related code that I first wrote, with Dick Bond and Lloyd Knox, about 30 years ago. This older algorithm (“BJK”, from our initials) answered the same statistical question as Almanac, using a very different technique. On large areas of sky, Almanac and this older algorithm got the same answer — the code works.

We went on a long rabbit-hole modifying the details of Almanac’s setup, making it more similar to other state of the art samplers. This change also didn’t solve our problem, although it seems to help on the margins. I made one suggestion that I thought would help, based on our long-ago experience with the BJK algorithm, bundling up some of the numbers we were trying to determine into “bands”. I don’t know if Claude would have come up with this idea on its own, and it took a while to get the details right. In fact, Claude would sometimes declare premature victory, admitting its mistakes only when I pointed them out.

It worked, eventually. After a lot of iterations, we transformed a problem unsolvable with the previous version of the code to one that was, well, easy.

But it only worked because my knowledge and experience — literally decades working on problems of this sort — meshed with Claude’s own “talents” — quick turnaround, patience, and encyclopaedic, if not always discriminating, knowledge of computing and of at least some aspects of the underlying science, statistics, and mathematics.

And it was fun! I thought that I liked programming, but I am very happy to have Claude do most of the grunt-work for me. The quick turnaround, and not having to sweat the details of writing and running re-writing and re-running program after program, was a delight.

In many way, working with Claude was like working with a junior colleague. But Claude is not a colleague, but a machine. And, as David Hogg has advocated, the point of doing astrophysics, a beautiful but useless field of science, is exactly the training and fulfilment of the people doing it. That has at least two implications. First, given how much my own experience was necessary to getting good results, that means we had better make sure that we are training humans, not just better LLMs. Second, no matter how delightful the interactions, they mustn’t replace training our students and collaborating with our colleagues.

(This post was written by me, not by Claude, though I did ask it to suggest a title — and this sentence.)

July 07, 2026

Jordan EllenbergUS World Cup watch party, Wingra Park

Big screen, big crowd to watch USMNT against Belgium. Team USA, sadly, didn’t do much to make it festive. The three North American host countries knocked out in three straight days. Sure was a nice night by the lake, though.

July 06, 2026

Doug NatelsonOMB proposed rule changes - act now

For non-US folks, feel free to skip. 

For US folks:  The Office of Management and Budget, which for much of its history has been a comparatively uncontroversial element of the executive branch, has set rules and guidelines for how many executive-branch agencies conduct business and interact with, e.g., universities.  For the purposes of how the research ecosystem operates, the most relevant is OMB's "Uniform Guidance" about how grants and contracts work.  Periodically these rules are updated for various reasons, including the goals and policies of the presidential administration.  The standard way this works is that the proposed changes are published in the Federal Register; there is a public comment period; OMB makes revisions and then publishes the new rules.  In principle, Congress can act to override or prevent rule changes, but without the agreement of the President, this is an extremely challenging path.

OMB has proposed sweeping changes to the Uniform Guidance, summarized here.  These proposed rule changes are huge deviations from previous practice.  For example, they would have all final grant decisions made by political appointees or hires of the executive branch (rather than, e.g., agency subject matter experts); grants could be cancelled at any time for essentially any reason (completely undefined insufficient support of the president's priorities), with no appeal process; international collaborations would be severely curtailed. That's just three for starters.  Note that this would also go beyond just the public research enterprise - it would allow the executive branch to cancel funding for things like bridges, roads, schools, agriculture, etc. for undefined political reasons.  It would be a huge transfer of power from Congress to the presidency.  Here is another summary by the AAU.  Here is an editorial essay from ars technica.  

The public comment period on this runs until July 13.  Here is a link where you can make a comment.  Here is a guide for how to be effective at this from Stand Up for Science.  The APS has a tool for helping people to comment about specific aspects of the rule changes.  It is also a good idea to contact congressional delegations (representatives, senators).  

It's important to have a clear public record about the proposed changes.  They may try to implement these regardless, but if so, there will be a continued fight over this in Congress and through the courts.  

July 05, 2026

Jordan EllenbergSave it for 2062, Peacock

I love the 1812 Overture as much as the next man, indeed substantially more, but I think it’s weird that Peacock’s “Star Spangled Sunday” baseball broadcast is celebrating America’s 250th birthday with a song about the Russian army repelling a French advance just short of Moscow.

Oh, hmm, apparently this is a thing.

Tommaso DorigoPawn Storm

Pawn Storm

I once was an active chessplayer, but work duties have long taken tournaments off my plate - I simply do not have the time to sit through long hours of chess battles. So I play blitz online on chess.com (my handle is "tommasodorigo", in case you wondered).

Tommaso Dorigo
Categories

Scott Aaronson Happy 250th!

I’m at the New Jersey shore with family and friends, where we’ve spent this Fourth of July eating hot dogs, playing miniature golf, and wading into the ocean that my great-grandparents crossed to escape calamities they knew about and much greater calamities that they didn’t. Tonight we’ll see the fireworks, weather permitting.

And yes, on the crowded beach today you can find people sporting MAGA and “45-47” hats, and even a giant “Trump 2028” flag—a stark reminder of the millions who would redefine the meaning of our 250-year-old experiment to something dark and authoritarian, the opposite of what its founders intended. Of course, those forces find mirror images on the left end of the political spectrum, where one can find millions more who fully agree with MAGA about the failures of liberalism and the Enlightenment, differing only on the secondary question of which racist thugs should rule instead.

Despite everything, I don’t believe that both factions together constitute a majority. Even on the beach, the MAGA hats are vastly outnumbered by “250” banners and girls in stars-and-stripes bikinis, Americans who just want to celebrate.

Despite everything, I remain thoroughly American if I’ve ever been anything, and invested in the country’s future if I’ve ever been invested in anything.

JD Vance and his friends, who might rule the country after the predictable failure of “Trump 2028,” make a huge deal about “Heritage Americans.” Of course the point of such phrases is to exclude those like me, and recent immigrants, and even (incredibly) JD’s own wife. On reflection, though: could I, too, count as a Heritage American at this point? After all, my family has now been here for half the country’s history. My grandfather, who grew up in poverty, became a professional boxer in Philadelphia and Atlantic City during the Great Depression. He then joined the Army and ended up clearing German mines in North Africa and Italy in WWII. He was assigned to a company of Southerners, who had never met a Jew and were shocked that my grandfather didn’t have horns—but by the war’s end, my grandfather and the relatively few others in his company who remained alive had become best friends. My grandfather told me that he could understand the German POWs who they captured tolerably well, since German was similar enough to Yiddish, but who he could never understand was the British.

As for me, I grew up in the town of Washington Crossing, PA, maybe a mile’s walk from where this happened (and where it’s still reenacted every Christmas):

My earliest childhood hero (that I can remember) was Ben Franklin, whose institute in Philadelphia I visited often. I didn’t even recognize as unusual at the time how the founding of the country didn’t feel like a remote abstraction to me, but was all around me, as if it was yesterday.

That the founders of the United States created the model for all time of how to bend the arc of human history a little bit away from its usual horribleness, of how to overthrow a despotism without instituting an even worse despotism in its place, of how to found a new civilization on ideas and principles rather than raw power … is one of those things that seemed true to me as a child and that still seems true to me today.

May this greatest experiment continue for another 250 years. May it triumph against all those within and without who would see it destroyed.

July 04, 2026

Tommaso DorigoWill The AI Bubble Burst?

Will The AI Bubble Burst?

As a forewarning to the reader, I am a physicist, not an economist. What I wrote below may very well be naive.

Tommaso Dorigo
Categories

July 03, 2026

Matt von HippelAmplitudes 2026

This week was Amplitudes, my old subfield’s big yearly conference. This year, it’s at Queen Mary University of London.

I’m too busy to attend these days, now that nobody is paying me a salary to do that sort of thing. But it’s still a good chance to keep up with the field, which helps me find stories. And I know I have a few readers who are interested. So I read the slides when I can, and fill you all in.

As of writing this, I’ve read through slides from the first few days of talks. I’ll likely post on the rest next week. As usual when I’m conference-blogging, this post will be quite a bit more technical than my average post, so readers beware: I’ll be mentioning a lot of amplitude-ish ideas without much explanation. I’m happy to explain in the comments if you’re curious, though!

Before I launch into talking about the content, I should mention something I’ve heard about the venue. Apparently this year registration filled up surprisingly quickly. The rumor is that the folks at Queen Mary weren’t able to find a venue with enough space to host the full community, leading to a smaller conference than usual. As the Amplitudes subfield grows, I suspect this will be more and more of a challenge. I remember back when I was organizing in 2021, rooms that could hold enough people were shockingly expensive to book. I wouldn’t be surprised if this becomes more of an issue going forward.

David Kosower opened the conference with a review of the state of the art in amplitudes for gravitational waves. I enjoyed an early slide showing how LIGO has increased in precision over the last ten years, which really helped illustrate the value that good theoretical predictions can bring as the experiment gets better. I also appreciated his attempt to get people to stop saying “post-Minkowskian” and start using a more normal term like Relativistic Perturbation Theory, though based on the other talks it doesn’t look like it’s catching on. After covering the overall state of the art (currently around four loops) he talked about his own work looking for ways to more directly get waveforms for orbiting black holes out of an amplitudes-style calculation, a theme his collaborator Donal O’Connell covered in more detail later that day. (The trick, apparently, is background fields!) Other gravitational wave-related talks came from Gustav Jakobsen, who covered four-loop results using the worldline method, Graham Brown, who explained how to calculate the Magnusian Jakobsen mentioned (it’s the log of the amplitude, basically), Canxin Shi, who talked about how a concept called Stratonovich-Weyl quantization helps explain structures that keep showing up in classical observables, Giulia Isabella, who covered a way to get six-loop contributions to gravitational waves via a wave equation, Lara Bohnenblust, who showed how to get strong-field information from a shockwave limit, and Gang Chen, whose slides were brief enough that I didn’t really get a clear idea of what he was up to. Rounding out the day were two talks that didn’t involve gravitational waves: Zihan Zhou, reporting on work with Nima on a water wave polytope called the hydrotope that they found a general formula for with help from Claude (the AI, not Duhr, to steal a recurring joke from Lancefest), and Michael Ruf, who probably snuck in due to the fact that he has done a lot of work with gravitational waves, but was reporting on a QCD calculation using a tool called Scorpio.

Tuesday had more of a QCD theme. Thomas Gehrmann gave the day’s opening review talk, where he pointed out that to get 1% precision for predictions for the LHC, we’ll likely need three-loop calculations. He pointed out the important role IR real emission calculations and improvements in parton distribution functions need to play, and pointed out that collider physics isn’t just the LHC: between reanalyses of older electron-positron collider data and the upcoming electron-ion collider and FCC, there are even more contexts where high-precision QCD will matter. Simone Zoia talked about one aspect of the current state of the art, five-particle processes involving massive particles, where the functions get strange and elliptical and one has to think hard about the methods one uses (for example, do you want canonical differential equations, or almost-canonical? Bulirsch-Stoer or AMFlow for numerics?) He also included an excellent Laplace quote, “Nature Laughs at the Difficulties of Integration”. Yang Zhang talked about a different calculational frontier, covering progress in the planar limit and with no masses, but for two-loop six-particle and three-loop five-particle scattering. The talk included some nice branding, like an “epsilon collaboration” proposing an algorithm to find epsilon forms for any Feynman integral and a program called Effortless to find symbol letters. Yang Zhang is also a skilled amateur photographer, and he mentions taking the conference photo for Amplitudes five times before. Personally, I’m surprised it’s only been five times! Dmitry Chicherin presented results bootstrapping QCD amplitudes. This was a dream of mine back in my hexagon function days, and while they aren’t quite at the point of being useful (my understanding is they’re still only getting the leading transcendentality, and none of the amplitudes they’re finding are new) it’s still pretty cool that this is even possible now. Bo Feng proposed what he claims is a general algorithm to find generating functions for IBP reduction. It’s not clear to me whether his setup bypasses the computational difficulties in existing methods (mostly involving solving large systems of equations), or whether it shifts the issue elsewhere, though. Pierre Vanhove’s talk also involved QCD, though at an effective field theory mediated step removed, with chiral perturbation theory, a theory of low-energy QCD that he used to shore up the lower energy ranges of lattice calculations for contributions to the muon anomalous magnetic moment (a natural calculation to involve him due to the presence of novel elliptic integrals). Overall, the QCD talks impressed me with the wide range of software packages mentioned, some of which already existed when I was in the field but many of which are new. I do wonder if this is just a result of many research paths maturing at the same time, or if people are finding it easier to write packages with tools like Claude Code.

The remaining three Tuesday talks were more mathematical or theoretical in focus. Henrik Johansson talked about black hole Compton scattering in N=8 supergravity, where it’s possible to use the wave equation to extract all-loop results. Cristian Vergu reported on his progress with Landau analysis. It’s been really fun watching this grow from a small reading group at NBI to what appears to be at this point quite a deep understanding, including a picture of how to understand amplitude singularities with quite a lot of breadth and detail, and the persistent hope that this could allow one to manipulate singular quantities without needing dimensional regularization. Michael Borinsky gave an update on tropicalization, where he showcased a theorem he proved demonstrating a way to compute certain classes of amplitudes in polynomial time in the loop order, even when the number of Feynman diagrams increases factorially. The examples he talked through were impressive, but I’m still a bit skeptical this could work for the Standard Model. I’d want to talk to him about it to figure it out, anyway!

Wednesday was a short day, as is tradition, to give people some time for tourism in the afternoon. Alexander Zhiboedov began the day with a talk on energy correlators in N=4 super Yang-Mills at finite coupling, where he is now able to do a bootstrap with two-sided bounds to hone in on the actual quantity even outside of the planar limit. Arthur Lipstein and Daniel Baumann both talked about cosmological correlators, the former with amplitudeologists’ favorite toy model of the conformally coupled scalar and the latter with Yang-Mills and gravity in de Sitter space. Dave Dunbar closed the day with a historical talk, walking through the milestone amplitudes papers before the Amplitudes conference existed. It’s the kind of talk he could have given at Lancefest the week before if others hadn’t already covered the material!

I’ll cover the rest of the conference (Thursday and Friday) in next week’s post, so see you then!

Scott Aaronson An American privacy emergency: Guest post from Cynthia Dwork et al.

Scott’s foreword: Cynthia Dwork is Gordon McKay Professor of Computer Science at Harvard, and a pioneer in the fields of differential privacy and algorithmic fairness. On my recent travels to the SigmaWest science camp and then STOC, there was much talk about a recent Trump administration action that would ban not only differential privacy, but essentially all modern techniques for preserving privacy in large datasets, for example in the 2030 US Census. I realize that many of us have “outrage fatigue,” but this particular outrage hits extremely close to home for the CS theory community. So when Cynthia approached me at STOC to propose a guest post on the issue, of course I said yes. The post that she sent me, below, is cosigned by many other leaders in the field.


On June 4, 2026, the U.S. Secretary of Commerce issued a directive (DAO 216-26) relegating confidentiality protection in all Bureau of Economic Analysis (BEA) and U.S. Census Bureau publications to techniques dating back to the early 1970s, turning its back on over half a century of progress and protections for data subjects. Advances in confidentiality provision had enabled the Census Bureau to share increasing quantities of data at more granular detail. The order will result in less useful (or fewer available) statistics, weaker protection, or both. We write to illustrate the danger posed by the order and to mobilize the scientific community to speak out against it.

The acting force behind this order is political interest, not scientific merit. DAO 216-26 bypassed legally required administrative procedures. It fulfills a promise made by the architects of the Heritage Foundation’s Project 2025, and reflects both the rhetoric and misunderstandings of representatives of the Center for Renewing America (CRA), an organization founded by OMB Director Russell Vought. CRA’s explainer on the use of differential privacy in the 2020 Census is up-front about the stakes: “Even if the citizenship question is added to the Census, it will be impossible to ascertain the status of individuals so long as differential privacy is used.” But masking this sort of personal characteristics data is legally required by the Census Act (13 U.S. Code Section 9), which makes it a crime to “make any publication whereby the data furnished by any particular [individual] can be identified.” Confidentiality is also widely understood as critical to ensuring that people respond to the census. 

DAO-216-26 bans differential privacy and other modern (and not so modern) techniques. It restricts disclosure avoidance techniques to “coarsening,” which it describes as “reducing the level of detail or specificity of published statistics, such as through rounding, aggregating (grouping), and/or the use of ranges.” “Suppression” (“expressly redacting certain values”) may also be used, but only as a “last resort.” DAO-216-26 forbids “noise infusion”, described as “methods that involve modifying a dataset by adding random values, or noise.” 

Noise infusion was invented precisely to address the increasing demand for granular data in the face of confidentiality laws that forbid publishing reidentifiable data. Coarsening and suppression were satisfactory for most national, aggregate statistical series, like the Principal Federal Economic Indicators. However, these techniques failed when applied to business and demographic data at fine geographic or industrial detail. By forbidding noise infusion, the directive bans the disclosure avoidance techniques at the core of dozens of data releases over the last three decades. It bans input noise infusion, used in the Quarterly Workforce Indicators since 2002 and, until now, planned for the Bureau of Economic Analysis statistics [1].  It bans swapping, used for decennial census publications since 1990. It also bans differential privacy,  the best currently known approach for obtaining the most data utility for any given level of privacy. Differential privacy was used for sharing data on commuting patterns (OnTheMap) since 2008 and for publications based on the 2020 Census. Until the recent directive, differential privacy was planned for the 2030 Census too.  Many other products and procedures are implicated as well.

1.      Illustrations

DAO-216-26 is incompatible with the Census Bureau’s dual mandate to provide confidentiality and fitness for use. To illustrate this, we recall and expand on an example due to Nathan Goldschlag, inspired by the County Business Patterns (CBP) data, which provides statistics on business activity broken down by industry and geography. Goldschlag describes three scenarios, illustrating the tension between providing useful information and maintaining confidentiality of responses as required by the Census Act. 

·       “There is only one brewery in a small county. If the CBP published the exact count of brewery employees in that county, it would be disclosing the information of one business (how many workers it employs), a clear violation of the law.2

·        “There are two breweries in a small county, and the CBP again publishes the exact count of brewery employees. If I own one of those breweries, I could learn how many employees my competitor has, again violating the law.

·        “There are more than two breweries in a small county, but the CBP chooses not to publish the total number of brewery employees out of concern that it might compromise the privacy of the businesses. If I’m a prospective brewery owner, I may deem the project too risky to pursue without information about the market I’m entering.”

In Goldschlag’s example, coarsening makes the published statistics useless.  We now add a fourth scenario, showing that it also fails to maintain confidentiality.  To keep things simple, assume none of us owns any of the businesses in the new example. The County has two towns with one brewery each, North Bend and South Bend. Furthermore, North Bend has a mobile bottling company and South Bend has a stationary bottling company. That’s a total of four beer-related business entities in the County.  Two of these businesses, the North-Bend brewery and the South Bend bottling company, are publicly-owned.

  • The CBP publishes five statistics:
    • (A) The total number of employees in beer-related businesses in North Bend: Because there is only one brewing company in North Bend and only one bottling company in North Bend, the category is coarsened to “beer-related”.
    • (B) The total number of employees in beer-related businesses in South Bend: Because there is only one brewing company in South Bend and only one bottling company in South Bend, the category is coarsened to “beer-related”.
    • (C) The total number of employees in brewing only:  Because there is only one brewing company in each of North Bend and South Bend, the statistic is coarsened to the total number of employees in brewing only in the County.
    • (D) The total number of employees in bottling only: Because there is only one bottling company in each of North Bend and South Bend, the statistic is coarsened to the total number of employees in bottling only in the County.
    • (E) The total number of employees at publicly owned companies: Because there is only one publicly owned company in each of North Bend and South Bend, the statistic is coarsened to the total number of employees in publicly owned companies in the County. 

We now have 5 equations in 4 unknowns. Using only 4 of these (A, B, C, and E), we can solve for the exact number of employees at each of the four companies with high school algebra.

In the above (fictional but realistic) scenario, the County Business Patterns were released with good-faith coarsenings for the geographical, business, and ownership categories.  Nonetheless, even without inside knowledge of one of the companies’ number of employees, we can completely reconstruct all four numbers.  What happened?  The coarsenings interacted poorly.  Noise infusion perturbs that set of equations, preventing exact reconstruction.

2.      Impediments to Implementation

The Commerce Department now claims the directive’s return to the outdated “tradstat” traditional statistical techniques of the 70s is good for data consumers: “This update to our disclosure limitation method protects respondents and provides the public with more essential economic information.”  (Emphasis added.)  As we saw from Goldschlag’s example, coarsening does just the opposite.

And it can’t be fixed. Coarsening by definition reduces access to fine-grained information.  Our example of three poorly interacting coarsenings shows that this sacrifice is for naught: without noise infusion, confidentiality is destroyed by elementary calculations.  For population surveys, this is precisely what formal noise infusion methods, like differential privacy, protect against; this is the “fancy math” that Goldschlag mentions in his post and that holds personal characteristics, like citizenship status, in confidence.  

3.      Confidentiality is Critical for Federal Statistics 

The scientific community continues to debate the best techniques for protecting the confidentiality of respondents’ data, but DAO-216-26 is not driven by science. It is driven by political interests. Those issuing this order are willing to risk the public’s trust in the process. We think that this is wrong-headed and dangerous. 

Civil servants will do their best to comply with this order while still following the laws that require them to protect the confidentiality of respondents’ data. To balance these competing mandates, they may seek to produce less data or coarsen data so much that it is unusable. Or they might be pushed by political actors to publish data that can be easily unmasked, like in the brewery examples above. Regardless of their choices, they will be hard-pressed to guarantee respondents’ confidentiality, which will prompt many businesses and individuals to simply not answer. This is devastating for an agency that delivers democracy’s data.

Conclusion

Rather than political actors overruling the government’s own statisticians, we need deep investment in our nation’s statistical agencies, ensuring that agencies have the staff and support to improve their methods using the best available tools. Regardless of how the scientific community feels about any specific privacy-enhancing technique, we must collectively reject this anti-scientific approach to governing federal statistics. Too much is at stake. 

How to Take Action

Share this post with others in your professional network and community.

Contact your Congressional representative and voice your concerns. Calling or writing to your representative is one of the most effective and easiest things a constituent can do that should only take a couple minutes of your time.

  1. Find your representative contact information here.
  2. State your concern. Here is a sample script: “My name is [Name], and I am a constituent from [City] in your district [ZIP CODE]. I am calling because I am concerned about the the U.S. Secretary of Commerce issued a directive (DAO 216-26) that wants to relegate confidentiality protection in all Bureau of Economic Analysis and U.S. Census Bureau data products and statistics to outdated and ineffective statistical techniques. If followed, this order will destroy the Commerce public data our nation relies on for important decisions, such as where to build necessary services for our community’s well-being. I want the DAO to be rescinded. I want proper administrative procedure to be followed. I want technical decisions such as the choice of method used to balance utility and confidentiality to be informed by professionals in the federal statistical agencies, not made unilaterally by political operatives.”
    1. Optional is stating what kind of constituent, such as a retired teacher or a working professional.

Volunteer to help preserve Census working papers and documentation. Pages explaining “noise infusion” and “differential privacy” are already going offline. Archive relevant methodology pages and technical documentation. You can also do this via the Internet Archive’s Wayback Machine (“Save Page Now”).

John Abowd 
Aloni Cohen
Cynthia Dwork
Jae June Lee
Jayshree Sarathy
Adam Smith
Salil Vadhan


[1] BEA Working Paper WP2026-9, now purged by the Department of Commerce.  As of 6/22 Google returns:

Bureau of Economic Analysis (BEA) (.gov)
https://bea.gov › files › papers › BEA-WP2026-9

July 01, 2026

Peter Rohde Aiguilles Crochues Traverse (2,840m)

An album of GoPro headcam footage from climbing the Aiguilles Crochues Traverse (PD, 2,840m) near Chamonix, France in 2022.

Full playlist (45 videos): https://youtube.com/playlist?list=PLM4i-DL0BZ0Q

June 30, 2026

Peter Rohde Frenchmans Cap (Sydney Route)

Footage from our climbing trip to Frenchmans Cap, Tasmania (Australian grade 17, 380m) in 2022.

Full playlist (47 videos): https://youtube.com/playlist?list=PLT0z6qQjCS3c

Peter Rohde Triglav, Slovenia (2,864m)

GoPro headcam footage from climbing Triglav (2,864m), highest mountain in Slovenia, via ferrata. Climbed in 2022.

Full playlist (40 videos): https://youtube.com/playlist?list=PLZgovD57Nsr4

June 29, 2026

John PreskillThe physicists of Florence

A scientist in Florence can’t avoid bumping into colleagues. 

When visiting the Renaissance’s birthplace last summer, I ran into a fellow physicist even on a Saturday morning. I was wandering around the Uffizi Gallery, a museum blessed with some of the greatest hits in western art. A familiar face arrested me on the first floor.

Another colleague cropped up outside the museum. (Some might classify him as an applied physicist or an engineer, but he exhibited a theoretical physicist’s overactive imagination.)

One colleague, I’d been looking forward to meeting for over four years. Jae Dong Noh is a professor of physics at the University of Seoul in South Korea. He’d conducted the first numerical tests (classical-computer simulations) of an idea I’d helped midwife, the non-Abelian eigenstate thermalization hypothesis (NAETH). An earlier blog post described this mouthful, which predicts how certain quantum many-particle systems thermalize, or experience the flow of time. These systems’ dynamics conserve properties, analogous to energy, that are incompatible: one can’t measure the properties simultaneously, as one can’t measure a quantum particle’s position and momentum simultaneously. Because incompatibility helps distinguish quantum from classical physics, such systems’ thermodynamics qualifies as particularly quantum.

Jae Dong modeled such a system and others numerically in a paper. I admired his computational techniques and his grasp of symmetries (for experts: how non-Abelian symmetries affect chaotic quantum systems’ energy-level statistics). My postdoc Aleks Lasek was planning a more thorough numerical test of the NAETH, so I reached out to Jae Dong, and a collaboration crystallized. 

Seoul operates thirteen hours ahead of Maryland, but we managed to Zoom because Jae Dong is a night owl and I’m an early bird.1 Zoom introduced me to a man perpetually dressed in a neat button-down shirt and sweater, silver overriding the black in his hair. The neatness extended to Jae Dong’s explanations: if Aleks and I didn’t understand one of his emails, he’d explain it quietly and calmly, untangling the confusion as though pulling a comb through wool.

The collaboration settled into a rhythm: I’d pose a question or propose a goal, Jae Dong would respond with an analytical calculation,2 I’d find holes in the calculation, Jae Dong would plug the holes, I’d re-check the argument’s logic, and we’d repeat the cycle. Had I been in Jae Dong’s shoes, I’d have swallowed the constant objections as I’ve swallowed grape-flavored cough medicine,3 but he always responded with equanimity—sometimes even good cheer—and a possible solution. Meanwhile, Aleks and then-undergraduate Jade LeSchack checked our analytical arguments numerically.

Florence flaunted a little steampunk during my visit.

So smoothly did the collaboration hum along that we coauthored two papers before ever meeting in person. One demonstrates numerically that two quantum many-body systems (for experts: nonintegrable Heisenberg models) obey the NAETH.4 In the other paper, we derive a symmetry relation from the NAETH. If the 17-syllable NAETH is a mouthful, the symmetry’s name is half a mouthful: a Kubo–Martin–Schwinger (KMS) relation. It’s important because (i) it enables us to calculate how rapidly a thermodynamic system responds to a stimulus, such as a weak magnetic field, and (ii) physicists go gaga over symmetries generally. 

The KMS relation constrains thermal states—essentially, systems that have temperatures. Your typical isolated many-particle quantum system looks thermal if you can observe just a small chunk of it at a time. Accordingly, Jae Dong and collaborators had proved that isolated many-particle quantum systems obey the KMS relation approximately. The larger the system, the more accurate the approximation. 

We extended his argument to systems whose dynamics conserve incompatible properties. Such an extension might sound simple, but its proof filled 24 pages of appendices. (For experts: Clebsch–Gordan coefficients are tricky blighters.) We discovered that, under certain conditions, incompatible conserved quantities can reduce the extent to which a quantum system obeys the KMS relation. Quantum incompatibility can augment deviations from conventional thermodynamics.

Italy’s architecture impressed me.

Jae Dong planned to present about our work at StatPhys, an international statistical-physics conference, which Florence was hosting in 2024. Throughout the two-and-a-half months before the conference, the KMS relation consumed our team. (For experts: Clebsch–Gordan coefficients are very tricky blighters.) I even hid in my hotel room, working and reworking our proofs, during another conference during that time. 

The toil paid off. We submitted our KMS manuscript for public scrutiny the day I flew to Florence—because not only Jae Dong would be representing our team at StatPhys. I was looking forward to meeting him there for the first time.

A corner of the hall where the StatPhys opening ceremony took place.

The StatPhys committee outdid itself. The opening ceremony unfolded in Florence’s Palazzo Vecchio, where members of the Medici dynasty once lived. Giorgio Parisi, who won a Nobel Prize for statistical physics in 2021, lectured at the ceremony.

Giorgio Parisi, with another history maker.

The meat of the conference took place in two other palaces, the Palazzo dei Congressi and the Palazzo degli Affari. In one of them, I met Jae Dong. Although we’d shown that quantum incompatibility can defy thermodynamic predictions, he met my expectations.

We discovered another thermodynamic phenomenon challenged by incompatible conserved quantities, so stay tuned for another paper and blog post. Some colleagues, one can’t avoid; others are worth engaging with again and again.

1 Aleks has confessed to night-owl habits, but physics motivates him to adapt. Some days, he’s emailed me results before even I’ve woken up. Who needs coffee when the thrill of discovery electrifies one minutes after one hops out of bed?

2 An exact calculation written out on paper, as opposed to a numerical, or approximate, calculation performed by a silicon-based classical computer.

3 Does anyone like the grape flavor? Why do companies bother producing it?

4 Rohit Patil and Marcos Rigol, too, have checked numerically that a system obeys the NAETH.

June 26, 2026

Doug NatelsonSome science/tech items - scrolls, nanostacks, and beyond

 Some brief science and technology items heading into the weekend:

  • IBM has reported making prototype chips for the "0.7 nm node".   As always, one should not interpret that size scale literally, since the effective diameter of a single silicon atom is around 0.2 nm.  The basic building block of their architecture here is the nanostack, which is a limiting case, somewhat 3D-integrated version of their nanosheet "gate all around" field effect transistors.  The fact that these structures can be made at this scale, reliably and en masse, is just phenomenal.  
     


  • I'd written previously about the Vesuvius Challenge, the attempt to use a combination of x-ray tomographic imaging and machine learning to read the carbonized ancient Roman scrolls found in a villa in Herculaneum, where they had been buried by the pyroclastic flow from the eruption in 79CE.  Well, they've managed to read a complete scroll - here's the preprint.  Very cool, and the hope is that among those scrolls might be books believed lost to history.
  • At the beginning of the month, Microsoft unveiled the next iteration of their approach to implementing topological qubits based on superconductor/semiconductor hybrid devices, as described here.  The relevant preprint is this one.  Some reporting on this is here.  This week, Nature published a comment on the prior work as well as the reply.  
  • There has been an explosion of research in recent years about trying to use electromagnetic cavities to tune the physical properties of condensed matter systems.  I'd discussed this here.  In the last couple of weeks, this preprint appeared, reporting that placing few-layer NbSe2 in an appropriate (THz) cavity can increase the superconducting transition temperature from 3.02 K to 3.41 K.  A 13% increase in \(T_{\mathrm{c}}\) is certainly interesting.
  • The incoming president of the National Academy of Sciences has a nice statement in Science.  The key passage for me:  "By its charter, the Academy is nonpartisan and neither a progressive organization nor a conservative one. It is a scientific body that follows the evidence wherever it leads, even when the destination might be unwelcome. In heated and polarized discourse, it is the Academy’s obligation to be the most careful and trustworthy voice. But rigorous science that arrives too late, or speaks too quietly, serves no one."  

Matt von HippelAt Lancefest

It’s been a while since I’ve said this: I’m at a conference this week!

Specifically, I’m at Lancefest, which is not just any old conference, but a birthday conference for Lance Dixon. When a renowned academic turns 60 or so, their students and collaborators hold a birthday party-flavored conference for them. The conferences are usually a mix of academic talks and reminiscences, with the occasional roast thrown in.

I went to my advisor’s birthday conference four years ago. Lance wasn’t my advisor, but in many ways he might as well have been. When my advisor took a sabbatical in the middle of my PhD, he sent me to work with Lance. It was my first real experience doing research in a team, not just puzzling away by myself with occasional feedback. And I was hooked: I spent the rest of my academic career in Lance’s field. We collaborated time after time, and even when I started to branch out he remained a frequent presence.

In part, that’s because Lance’s field was really Lance’s field. Amplitudeology has grown a lot since I started out, with several hundred people going to the field’s big yearly conference and subfields like Elliptics having their own yearly conferences. In such a world, it’s tough for anyone to feel like a truly central figure. But Lance tends to. He’s been able to keep up with that growing world, to keep finding important problems and keep understanding others’ ideas. While some have specialized, or stepped back, Lance seems to somehow manage to be a father figure for the whole field all at once. It’s a capability his advisor Jeff Harvey might have predicted, he mentioned in his talk that Lance’s interests were always broad. Over the years the young students I met who joined when the amplitudes field was already large saw Lance as a kind of mysterious titan, and were occasionally awed that I had worked with him. “What was that like?”

Well, it was like working with Lance. Lance isn’t a manager at heart, like some senior academics end up. He wants to understand everything he works on, and will happily dive in, Maple subscription in hand, and try to figure things out for himself. He wants you to keep up with him, understanding on your own terms, to keep him honest, to provide an independent check. But that often wasn’t possible, because the man is just so damn fast. He’d be miles ahead of me, with his Maple and his laptop, while I was churning overnight Mathematica runs on a twenty-machine cluster.

(Was the difference due to our taste in software? Partly. I did learn to use Maple later, it genuinely is faster at some things. But Lance is faster at almost everything.)

And he cared so damn much sometimes. About getting things right, like a scientist should. About making things nice, too: finding a pretty basis of functions, a better notation for the paper, something that might jostle out the next big insight. Working with him, you could feel like that one paper was the most important thing in the universe.

Others at the event have had similar stories. Fernando Febres Cordero remembers noticing a potential issue, emailing Lance about it, and in a few minutes hearing back with a potential explanation.

Lance is someone who became a leader without really being a politician. He doesn’t have the legions of students in tenured positions that some do. I trimmed that count by one, and it wasn’t huge to begin with. But for someone who isn’t “everywhere” in that sense, he manages to be “everywhere” all the same.

So Lance, happy birthday! You’re really the only person who could have had a birthday conference quite like this, a cross-section of the field, all with something kind to say. Thanks for putting up with any embarrassment associated with having this much attention for three days, and I wish you many Maple-fueled mysteries to come.

Peter Rohde Monte Rosa Tour — Spaghetti Tour (4,554m)

An album of one minute GoPro headcam clips from the Monte Rosa Tour (aka Spaghetti Tour) along the Swiss/Italian border in the Alps. This footage is from our expedition in 2022.

Full video playlist (101 clips): https://youtube.com/playlist?list=PLeRYUHyTyGxYa4zZnUIfiAScN1Hcl_HWe

June 23, 2026

Doug NatelsonBad to worse at NSF? (June 2026 edition)

This is a US research ecosystem post.  Feel free to skip if this isn't your cup of tea.

I'm showing my internet age in thinking that the right image to put at the top of this post is either the Drudge Report emergency light icon or an animated Star Trek "red alert" sign.

As you may be aware, NSF spending is incredibly low this year.  How low is "low"?  Check out this graph from grant-witness.  


This is a funding trajectory that has not been seen since the 1970s.  Now, because of budget uncertainties and disruptions last year, there was a big burst of activity late in FY25, and eventually the NSF did end up spending about what it was budgeted.  I spoke with one program officer at NSF last month who said that they fully intended to get there again this year, even if it meant he didn't have a vacation until September.  

A lot of people had looked at the trajectory above and worried that we are headed toward some kind of very bad outcome.  For example, if NSF is underspent by $3B by August, whether because of direct OMB opposition or because the award office at NSF is told by political leadership not to make awards, then it might be nearly impossible for NSF to spend its budget, at which point there could be a pocket rescission.  Basically, the executive branch has wanted to enact 50+% cuts to NSF; Congressional appropriators have said "no", but the executive branch may be trying to get their cuts anyway.  This would a terrible precedent.  If it happened you might expect Congress to be upset that their appropriations were being ignored.  There would likely be lawsuits.

Today, however, this story broke in Science.  Supposedly, there are going to be broad cuts to many parts of the NSF, at the level of 20-30% in the present fiscal year, despite the fact that the NSF budget is only down 3% from last year and there is statutory language in the appropriation bill saying that no directorate could be cut by more than 5%.  

The article basically says that it is likely that the funds are going to support the X-Labs effort run out of the TIP directorate.  What is an X-Lab?  I have some inkling because I attended the webinar about the present solicitation a couple of weeks ago.

The idea of X-Labs comes from proposals like this.  The basic premise is (1) The present system holds back innovation for some and we need to be more flexible and entrepreneurial.  (2) We could bring together teams of people who could be in a position to do something transformative, with definite technology applications, but whose work is at an early stage such that it's too low a technology readiness level to attract VC/angel investors who could support a startup, or is too far off from deployment to be partnered with industry as in the long running SBIR/STTR program.  Thus, this team of people would form an X-Lab, where the key investment (say $50M/yr for 3-5 years, in a milestone-driven contracting method) would come from NSF/TIP.   This is not a priori crazy - multiple other groups have looked at non-profit startups as a way to fund science.  The program was announced at a level of $150M/yr for ten years.  (The Science article implies that those in charge want a lot more money now than was in the TIPS appropriation plan for this year.  Here is a claim that this is not true, which would make the cuts even harder to understand.)

One big catch:  The way the X-Labs are being implemented seems pretty inflexible.  An X-Lab has to be its own entirely independent ("autonomous") entity (rather like a company or non-profit), not a subsidiary or an operating unit of a company or university.  Any senior personnel involved are required to be 100% full-time associated with the X-Lab.  That means that anyone doing this from a company or national lab would have to quit their previous job or go on a complete leave of some kind.  Anyone doing this from a university would have to resign their faculty position or go on a complete leave of some kind.  Issues like IP and benefits/health insurance seem nontrivial and not worked out.  Given the current uncertainties with everything associated with the NSF, this is quite a proposition for established researchers to undertake.  

So, here we are, with reporting that there will be large cuts across the NSF, regardless of what the appropriations said.  Anyone with first-hand knowledge who wants to chime in, please weigh in in the comments, or drop me a line (presumably from a non-NSF email address).  

As bad as this is, the part of the article that truly angered me was this:
Program managers would normally rush to inform potential and current grantees about such dramatic changes. But the memo tells program managers to keep their mouths shut. “This information is highly confidential,” it reads. “Please do not communicate anything to PIs [principal investigators].”
Really?

You know this is not supported by the actual program officers, because this "highly confidential" information was almost immediately sent to a reporter.  Daylight is a great disinfectant.  Public pressure and Congressional pushing forced NSF leadership to relent on the plan to destroy the Ocean Observatories Initiative.  Maybe making this budget cutting known can focus attention on this, rather than having drastically reduced NSF research funding be a fait accompli.


June 20, 2026

Tommaso DorigoA Visit to the Network School in Forest City

A Visit to the Network School in Forest City

Last week I traveled to Singapore to give an invited talk at the AI4X conference, an exciting new venue gathering computer scientists, physicists and engineers to discuss how AI will accelerate scientific discoveries. My talk was the second of the opening day in the plenary session.

Tommaso Dorigo

June 19, 2026

Matt von HippelRadiation Radiates

I recently finished reading The Orphan Master’s Son, a (Pulitzer-winning, apparently) novel set in 2000’s-era North Korea. In one plot point, Kim Jong Il has agents steal a Japanese telescope designed to measure the cosmic microwave background radiation, under the mistaken impression that it will help him find uranium.

The novel plays it for (horrified) laughs, but I’ve seen this kind of misunderstanding crop up in the real world too. Sure, most people would realize that a telescope probably won’t help you find something buried under a mountain of rock. But there’s a deeper misunderstanding here. Ask yourself: what does “radiation” mean?

We talk about radioactive elements like uranium releasing radiation. We talk about electromagnetic radiation, including everything from gamma rays to visible light to the 5G of your cell phone. We talk about cosmic radiation coming in from space, and about the cosmic background radiation that originated in the early universe. For someone who doesn’t know much about physics, it probably sounds like all of these are the same kind of thing.

But they’re not!

It’s helpful to break things down in terms of particles. Radioactive elements release three main types of radiation: alpha, beta, and gamma. Alpha radiation consists of helium nuclei: two protons stuck together with two neutrons. Beta radiation consists of electrons. Gamma radiation is a type of electromagnetic radiation, and consists of photons: particles of light.

Anything we call electromagnetic radiation is a wave in the electromagnetic field, a ripple that moves through space. That’s different from other shapes of electromagnetic fields, like a magnetic field that stays in place. From a particle perspective, an electromagnetic wave is made up of photons, and physicists will often describe all such waves as light. Some of that light is the familiar rainbow of visible light, while some has lower-energy photons, like microwaves and radio waves, or higher-energy photons, like gamma rays or X-rays.

Cosmic radiation (more often called cosmic rays), like radiation from radioactive elements, can be many types of particles again. Most of it consists of protons, while some consist of various nuclei, or electrons. A smaller fraction are antimatter, like antiprotons or positrons. Sometimes, physicists include neutrinos when they talk about cosmic rays, while sometimes they include gamma rays.

The cosmic background radiation is once again different. This is an overall hum of microwaves, electromagnetic radiation from the early universe that has gotten fainter and more diffuse over time. Cosmologists will sometimes talk about when the universe was “radiation-dominated” versus “matter-dominated”. They’re referring to times when most of the energy of the universe was in electromagnetic radiation, versus when it was mostly in other particles.

The only thing that ties all of these meanings together is the word’s literal meaning: radiation radiates. It starts in one place and travels outwards, having an effect at a distance. For the first scientists to observe phenomena like X-rays, this was almost all they knew about them, so they tossed them together in one category. Now, we know much more, but the names stuck.

So if you hear a physicist use the word “radiation”, try to avoid making any assumptions. You can’t know, just from that word, what they mean.

And please, don’t steal any Japanese space telescopes.

June 16, 2026

n-Category Café Octonions and the Standard Model (Part 14)

Paul Schwahn and I have come out with a new paper about octonions and the Standard Model:

It builds on things I’ve discussed here, but it goes further. Let me explain a bit.

A bit is just a binary alternative: 1 or 0, true or false. That’s how it works in classical logic. We could also have a ‘trit’, meaning 3 alternatives.

In quantum physics we instead have qubits and qutrits.

Qubits and qutrits are usually described using complex numbers. The algebra of observables of a qubit is the Jordan algebra 𝔥 2()\mathfrak{h}_2(\mathbb{C}), consisting of 2×22 \times 2 self-adjoint complex matrices. Similarly, the algebra of observables of an qutrit is the Jordan algebra 𝔥 3()\mathfrak{h}_3(\mathbb{C}), consisting of 3×33 \times 3 self-adjoint complex matrices.

We can also study systems with more than 3 alternative ways to be. They work the same way, using the Jordan algebras 𝔥 n()\mathfrak{h}_n(\mathbb{C}) with n>3.n \gt 3.

But we can also do quantum mechanics using other number systems! The options have been mapped out, and the largest allowed number system for this purpose is the algebra of octonions.

A weird thing is that Jordan algebras built using octonions can describe qutrits, but not quantum systems with more than 3 alternative ways to be. The algebra of observables of an octonionic qutrit is the so-called ‘exceptional’ Jordan algebra 𝔥 3(𝕆)\mathfrak{h}_3(\mathbb{O}), consisting of 3×33 \times 3 self-adjoint octonion matrices. What makes it exceptional is that 𝔥 n(𝕆)\mathfrak{h}_n(\mathbb{O}) is not a Jordan algebra when nn is bigger than 3.

So, there’s something special about octonionic qutrits — and it turns out that every symmetry in the gauge group of the Standard Model is a symmetry of an octonionic qutrit!

Not every symmetry of an octonionic qutrit is a symmetry of the Standard Model. But those that do have a simple description. They are those that restrict to give symmetries of an ordinary qutrit sitting inside the octonionic qutrit… and an ordinary qubit sitting inside that!

That sounds exciting, but also vague, so let me make it precise.

While lots of people say the gauge group of the Standard Model of particle physics is U(1)×SU(2)×SU(3)\text{U}(1) \times \text{SU}(2) \times \text{SU}(3), in fact a certain subgroup of this acts trivially on all known particles. If we mod out by that, we’re left with

S(U(2)×U(3)) = {xSU(5):x=(* * 0 0 0 * * 0 0 0 0 0 * * * 0 0 * * * 0 0 * * *)}. \begin{array}{ccl} \text{S}(\text{U}(2) \times \text{U}(3)) &= & \Big\{ x \in \text{SU}(5) : x = \left( \begin{array}{c c c c c} \ast & \ast & 0 & 0 & 0 \\ \ast & \ast & 0 & 0 & 0 \\ 0 & 0 & \ast & \ast & \ast \\ 0 & 0 & \ast & \ast & \ast \\ 0 & 0 & \ast & \ast & \ast \end{array} \right) \; \Big\}. \end{array}

and this is the group I’m talking about.

We proved two theorems describing this group in terms of the symmetries of an octonionic qutrit. The group of automorphisms of the exceptional Jordan algebra 𝔥 3(𝕆)\mathfrak{h}_3(\mathbb{O}) is a 52-dimensional Lie group known affectionately as F 4\text{F}_4 — so that’s what I mean by the symmetries of an octonionic qutrit.

Here’s our main result:

Theorem 1. Suppose X,BX,B are Jordan subalgebras of 𝔥 3(𝕆)\mathfrak{h}_3(\mathbb{O}) such that

X𝔥 2(),B𝔥 3(),XB. X \cong \mathfrak{h}_2(\mathbb{C}), \;\; B \cong \mathfrak{h}_3(\mathbb{C}), \;\; X \subset B.

Then

Stab(X)Stab(B) 0S(U(2)×U(3)). \text{Stab}(X) \cap \text{Stab}(B)_0 \cong \text{S}(\text{U}(2) \times \text{U}(3)).

Here Stab(X)\text{Stab}(X) is the stabilizer of XX — that is, the subgroup of F 4\text{F}_4 consisting of elements that map XX to itself — while Stab(B) 0\text{Stab}(B)_0 is the identity component of the stabilizer of BB.

This ‘identity component’ business is rather sneaky, but it turns out that guys in Stab(B) 0\text{Stab}(B)_0 are symmetries of an ordinary qutrit that can be described as unitary operators on \mathbb{C}, while Stab(B)\text{Stab}(B) also contains those symmetries that are described by antiunitary operators. The CPT symmetry of the Standard Model is antiunitary, for example.

Theorem 1 emerged from a related result, which grew out of the work of Todorov and Dubois-Violette:

Theorem 2. Suppose A,BA,B are Jordan subalgebras of 𝔥 3(𝕆)\mathfrak{h}_3(\mathbb{O}) such that

A𝔥 2(𝕆),B𝔥 3(),AB𝔥 2(). A \cong \mathfrak{h}_2(\mathbb{O}), \;\; B \cong \mathfrak{h}_3(\mathbb{C}), \;\; A \cap B \cong \mathfrak{h}_2(\mathbb{C}).

Then

Stab(A)Stab(B) 0S(U(2)×U(3)). \text{Stab}(A) \cap \text{Stab}(B)_0 \cong \text{S}(\text{U}(2) \times \text{U}(3)).

Todorov and Dubois–Violette proved this for a certain standard choice of subalgebras AA and BB. Thus, the challenge in proving Theorem 2 was to show that every other choice can be mapped to this standard choice using the action of F 4\text{F}_4. This shows that the theorem is not an artifact of a specific choice, but rather a general fact.

How do we prove these results?

We start by constructing the octonion product from SU(3)\text{SU}(3)-invariant operations on \mathbb{C} and 3\mathbb{C}^3. We then use this description to reprove Todorov and Dubois–Violette’s special case of Theorem 2. Then we show that F 4\text{F}_4 acts transitively on the set of subalgebras of 𝔥 3(𝕆)\mathfrak{h}_3(\mathbb{O}) that are isomorphic to 𝔥 3()\mathfrak{h}_3(\mathbb{C}). We also show every Jordan subalgebra of 𝔥 3(𝕆)\mathfrak{h}_3(\mathbb{O}) isomorphic to 𝔥 2()\mathfrak{h}_2(\mathbb{C}) is contained in a unique Jordan subalgebra isomorphic to 𝔥 2(𝕆)\mathfrak{h}_2(\mathbb{O}). This lets us prove that F 4\text{F}_4 acts transitively on the set of pairs of Jordan subalgebra A,B𝔥 3(𝕆)A, B \subset \mathfrak{h}_3(\mathbb{O}) with A𝔥 2(𝕆)A \cong \mathfrak{h}_2(\mathbb{O}), B𝔥 3()B \cong \mathfrak{h}_3(\mathbb{C}) and AB𝔥 3()A \cap B \cong \mathfrak{h}_3(\mathbb{C}). Theorem 2 then follows from Todorov and Dubois-Violette’s special case. We conclude by using these results to prove Theorem 1.

However, if you want to get into the details of the physics, the interesting part is how the strong force gauge group SU(3)\text{SU}(3) and the electroweak S(U(1)×U(2))\text{S}(\text{U}(1) \times \text{U}(2)) show up from the relation between octonionic qutrits, complex qutrits and complex qubits. You’ll see that in the proof of Lemma 4.

And if you want to get into the details of the math, the main interesting thing here is the use of Jordan algebra technology like ‘Peirce decompositions’ and ‘Jordan frames’ to figure out what it must be like when you have a Jordan algebra 𝔥 2(𝕃)\mathfrak{h}_2(\mathbb{L}) or 𝔥 3(𝕃)\mathfrak{h}_3(\mathbb{L}) sitting inside 𝔥 3(𝕂)\mathfrak{h}_3(\mathbb{K}), where 𝕃\mathbb{L} is some normed division algebra contained in a bigger normed division algebra 𝕂\mathbb{K}.

What it all ‘really means’, if anything, is a question for later. It could be just a coincidence. Of course I hope not.

John BaezOctonions and the Standard Model

Paul Schwahn and I have come out with a new paper about octonions and the Standard Model:

The Standard Model gauge group from the exceptional Jordan algebra

It builds on things I’ve discussed here, but it goes further. Let me explain a bit.

A bit is just a binary alternative: 1 or 0, true or false. That’s how it works in classical logic. We could also have a ‘trit’, meaning 3 alternatives.

In quantum physics we instead have qubits and qutrits.

Qubits and qutrits are usually described using complex numbers. The algebra of observables of a qubit is the Jordan algebra \mathfrak{h}_2(\mathbb{C}), consisting of 2 \times 2 self-adjoint complex matrices. Similarly, the algebra of observables of an qutrit is the Jordan algebra \mathfrak{h}_3(\mathbb{C}), consisting of 3 \times 3 self-adjoint complex matrices.

We can also study systems with more than 3 alternative ways to be. They work the same way, using the Jordan algebras \mathfrak{h}_n(\mathbb{C}) with n > 3.

But we can also do quantum mechanics using other number systems! The options have been mapped out, and the largest allowed number system for this purpose is the algebra of octonions.

A weird thing is that Jordan algebras built using octonions can describe qutrits, but not quantum systems with more than 3 alternative ways to be. The algebra of observables of an octonionic qutrit is the so-called ‘exceptional’ Jordan algebra \mathfrak{h}_3(\mathbb{O}), consisting of 3 \times 3 self-adjoint octonion matrices. What makes it exceptional is that \mathfrak{h}_n(\mathbb{O}) is not a Jordan algebra when n is bigger than 3.

So, there’s something special about octonionic qutrits—and it turns out that every symmetry in the gauge group of the Standard Model is a symmetry of an octonionic qutrit!

Not every symmetry of an octonionic qutrit is a symmetry of the Standard Model. But those that do have a simple description. They are those that restrict to give symmetries of an ordinary qutrit sitting inside the octonionic qutrit… and an ordinary qubit sitting inside that!

That sounds exciting, but also vague, so let me make it precise.

While lots of people say the gauge group of the Standard Model of particle physics is \text{U}(1) \times \text{SU}(2) \times \text{SU}(3), in fact a certain subgroup of this acts trivially on all known particles. If we mod out by that, we’re left with a group called \text{S}(\text{U}(2) \times \text{U}(3)), which is

\Big\{ x \in \text{SU}(5) : x =   \left(   \begin{array}{c c c c c}  \ast & \ast & 0 & 0 & 0 \\  \ast & \ast & 0 & 0 & 0 \\  0 & 0 & \ast & \ast & \ast \\  0 & 0 & \ast & \ast & \ast \\  0 & 0 & \ast & \ast & \ast   \end{array}  \right) \; \Big\}.

and this is the group I’m talking about.

We proved two theorems describing this group in terms of the symmetries of an octonionic qutrit. The group of automorphisms of the exceptional Jordan algebra \mathfrak{h}_3(\mathbb{O}) is a 52-dimensional Lie group known affectionately as \text{F}_4—so that’s what I mean by the symmetries of an octonionic qutrit.

Here’s our main result:

Theorem 1. Suppose X,B are Jordan subalgebras of \mathfrak{h}_3(\mathbb{O}) such that

X \cong \mathfrak{h}_2(\mathbb{C}), \;\; B \cong \mathfrak{h}_3(\mathbb{C}), \;\; X \subset B.

Then

\text{Stab}(X) \cap \text{Stab}(B)_0 \cong \text{S}(\text{U}(2) \times \text{U}(3)).

Here \text{Stab}(X) is the stabilizer of X—that is, the subgroup of \text{F}_4 consisting of elements that map X to itself—while \text{Stab}(B)_0 is the identity component of the stabilizer of B.

This ‘identity component’ business is rather sneaky, but it turns out that guys in \text{Stab}(B)_0 are symmetries of an ordinary qutrit that can be described as unitary operators on \mathbb{C}, while \text{Stab}(B) also contains those symmetries that are described by antiunitary operators. The CPT symmetry of the Standard Model is antiunitary, for example.

Theorem 1 emerged from a related result, which grew out of the work of Todorov and Dubois-Violette:

Theorem 2. Suppose A,B are Jordan subalgebras of \mathfrak{h}_3(\mathbb{O}) such that

A \cong \mathfrak{h}_2(\mathbb{O}), \;\; B \cong \mathfrak{h}_3(\mathbb{C}), \;\; A \cap B \cong \mathfrak{h}_2(\mathbb{C}).

Then

\text{Stab}(A) \cap \text{Stab}(B)_0 \cong \text{S}(\text{U}(2) \times \text{U}(3)).

Todorov and Dubois–Violette proved this for a certain standard choice of subalgebras A and B. Thus, the challenge in proving Theorem 2 was to show that every other choice can be mapped to this standard choice using the action of \text{F}_4. This shows that the theorem is not an artifact of a specific choice, but rather a general fact.

How do we prove these results?

We start by constructing the octonion product from \text{SU}(3)-invariant operations on \mathbb{C} and \mathbb{C}^3. We then use this description to reprove Todorov and Dubois–Violette’s special case of Theorem 2. Then we show that \text{F}_4 acts transitively on the set of subalgebras of \mathfrak{h}_3(\mathbb{O}) that are isomorphic to \mathfrak{h}_3(\mathbb{C}). We also show every Jordan subalgebra of \mathfrak{h}_3(\mathbb{O}) isomorphic to \mathfrak{h}_2(\mathbb{C}) is contained in a unique Jordan subalgebra isomorphic to \mathfrak{h}_2(\mathbb{O}). This lets us prove that \text{F}_4 acts transitively on the set of pairs of Jordan subalgebra A, B \subset \mathfrak{h}_3(\mathbb{O}) with A \cong \mathfrak{h}_2(\mathbb{O}), B \cong \mathfrak{h}_3(\mathbb{C}) and A \cap B \cong \mathfrak{h}_3(\mathbb{C}). Theorem 2 then follows from Todorov and Dubois-Violette’s special case. We conclude by using these results to prove Theorem 1.

However, if you want to get into the details of the physics, the interesting part is how the strong force gauge group \text{SU}(3) and the electroweak \text{S}(\text{U}(1) \times \text{U}(2)) show up from the relation between octonionic qutrits, complex qutrits and complex qubits. You’ll see that in the proof of Lemma 4.

And if you want to get into the details of the math, the main interesting thing here is the use of Jordan algebra technology like ‘Peirce decompositions’ and ‘Jordan frames’ to figure out what it must be like when you have a Jordan algebra \mathfrak{h}_2(\mathbb{L}) or \mathfrak{h}_3(\mathbb{L}) sitting inside \mathfrak{h}_3(\mathbb{K}), where \mathbb{L} is some normed division algebra contained in a bigger normed division algebra \mathbb{K}.

What it all ‘really means’, if anything, is a question for later. It could be just a coincidence. Of course I hope not.

June 06, 2026

n-Category Café A New Blog

Readers may have noticed that I haven’t been very active here for a while. That isn’t because I haven’t felt the “blogging urge”, but because I felt that the things I want to blog about right now wouldn’t be very interesting to much of the n-Category Cafe audience: they’re mostly fairly technical details about implementing proof assistants (because that’s what I’m mostly working on right now).

Accordingly, I’ve started a new blog! It’s at https://gwaithimirdain.github.io/blog/. (Gwaith-i-Mírdain is the github organization for development of Narya, the experimental proof assistant for Higher Observational Type Theory – and now Multimodal Type Theory as well – that I’ve been spending most of my time on, and will primarily be blogging about.) And I already wrote three posts (mostly about implementing multimodal type theory, with several survey questions for the reader), so you can check it out right now and see whether it’s likely to be your cup of tea.

Never fear, I’ll still come back here when I have more category-theoretic things to write about.

June 02, 2026

John BaezInterview with Micah Zarin

I’m not completely happy with this interview with Micah Zarin. It was nothing he did, it was me. I forgot to say that current-day AI wastes a lot of energy, and companies hope to use it to lay off people, and oligarchs are using it to extract lots of money from everyone. While obvious, these things are tremendously important and I should have emphasized them.

I was distracted by Micah’s fear that AI would make a career in math pointless, which really surprised me. So instead of giving my general thoughts on AI, I focused on putting myself in his place and imagining what to do in that situation. I suggested doing math with the help of AI as a way to overcome his fear and go ahead doing math while keeping abreast of new developments. If AI overtakes humans in math in his lifetime, which is far from certain, this could be a way to keep productively participating in math throughout this process. But I warned him to be very critical of what LLMs say, to lessen the danger of getting caught up in the ‘AI vortex’ that is turning many people into crackpots.

Mathematics, in case you haven’t been paying attention, is different from some other subjects because it’s an area where LLMs have shown some truly impressive problem-solving ability: read the various mathematicians’ comments in Remarks on the disproof of the unit distance conjecture where they grapple with this. But nobody really knows where this is going. So far LLMs have not shown much ability to invent new theories of mathematics, so it would be jumping to conclusions to assume AI will soon overtake humans in that realm. It would also be jumping to conclusions to assume it won’t.

Whatever happens, the real danger is not that AI will become too good, but that it will become too evil—most likely because of the oligarchs, corporations and governments behind it. I wish I had emphasized that point, which is always on my mind.

I think I succeeded in making another point, which is that life will not become pointless simply because some other entity gets better than humans at something and knocks us off our throne. To think that the meaning of life resides in our superiority is a childish attitude.

John BaezSumming the Reciprocals of Primes

The sum of the reciprocals of the primes diverges, but very slowly. The sum of the reciprocals of the first 100 primes is

2.106…

The sum of the reciprocals of the first 1,000 primes is

2.457…

For the first 10,000 it’s

2.709…

And it keeps creeping up, ever more slowly. To get the sum to reach 6, you need to add up the reciprocals of the first 3 × 10¹³² primes—far more than the number of atoms in the observable universe! Luckily there is no shortage of primes.

Here’s how you can see that the sum diverges, and that it diverges very slowly. First, remember Euler’s product formula for the Riemann zeta function:

\displaystyle{ \sum_{n=1}^\infty \frac{1}{n^{s}} = \prod_{p \text{\, prime}}\left(1-\frac{1}{p^{s}}\right)^{-1} }

This diverges logarithmically when s = 1. Taking logs we see

\displaystyle{ \sum_{p \text{\, prime}}\ln\!\left(1-\frac{1}{p}\right)^{-1} \;\approx\; \sum_{p \text{\, prime}}\frac{1}{p} }

must diverge too—but only log-logarithmically!

A deeper result, called Merten’s Second Theorem and proved here, says that

\displaystyle{ \sum_{p \le n} \frac{1}{p} = \ln\ln n + M + o(1) }

for some constant M. This constant is called the Meissel–Mertens constant. You can think of it as a fancier relative of Euler’s constant \gamma, which is defined by

\displaystyle{   \sum_{k = 1}^n \frac{1}{k} = \ln n + \gamma + o(1) }

But it’s much less widespread in mathematics than Euler’s constant, much as primes are less widespread than natural numbers.

Here’s how it works:

What’s a bit surprising, given the slowness of convergence and the somewhat erratic behavior of the primes, is that people can compute the Meissel–Mertens constant very precisely:

M \approx 0.26149721284764278375542683860869585905\ldots

The trick, of course, is to use another formula for this constant, which lets you compute it much more efficiently than the definition. Here it is:

\displaystyle{ M = \gamma + \sum_{n = 2}^\infty \frac{\mu(n)}{n} \ln \zeta(n)}

where \mu is the Möbius function and \zeta is the Riemann zeta function. By the way, this formula shows that calling M a fancier relative of \gamma is not just talk.

When you know how to compute the Meissel–Mertens constant, you can compare the sum of reciprocals of primes to \ln\ln n + M, and the agreement is very good:

In 1983, Guy Robin proved the curve goes above and below the actual sum infinitely many times, i.e.

\displaystyle{\left(\sum _{p\leq n}{\frac {1}{p}}\right) -\ln \ln n-M}

changes sign infinitely many times. You can see a bit of that happening here:

Puzzle. Can you find a naturally occurring sum that diverges even more slowly, for example like \ln \ln \ln n ?

Acknowledgements

The pictures were created by Dcoetzee, Marek Wolf and Saroad, respectively, and placed into the public domain on Wikicommons. Click on the pictures for more details.

May 28, 2026

John PreskillUnleashing the Advantage of Quantum AI

As experimental capabilities advance rapidly, the quantum computing community faces a critical elephant in the room: What will these quantum machines eventually be useful for? Will they deliver the promised broad societal impact, or will they remain highly specialized devices for exotic tasks known only to the experts?

The elephant in the room

Despite decades of effort, conclusive evidence of large quantum advantage in real-world applications remains confined to a few niche domains, such as simulating quantum materials and cryptanalysis. These problems are either inherently quantum to begin with, or they possess specialized mathematical structure that quantum algorithms can easily exploit. But it seems unlikely that such structures appear broadly in everyday life.

Indeed, most applications of modern computation hinge on the processing of massive, noisy classical data, generated at an unprecedented pace across society. That is the driving force behind the overwhelming success of machine learning and AI. Since the data originates from the macroscopic classical world, there is no obvious reason it should exhibit the delicate, specialized structures that quantum computers require. To playfully adapt Richard Feynman’s famous quote: We live in an effectively classical world, dammit, and maybe classical computers and AI already suffice for most of our problems. (For those unfamiliar, Feynman originally quipped: “Nature isn’t classical, dammit, and if you want to make a simulation of nature, you’d better make it quantum mechanical.”)

The central challenge

To truly unlock the power of a quantum computer, quantum algorithms typically need to access data in quantum superposition, processing many different samples simultaneously in different branches of the quantum multiverse. To use technical jargon, this is called querying a quantum oracle. But in reality, the classical data samples that we want to process are generated from everyday activities in a classical world, and we can only access them one at a time.

Think of the movie reviews you scroll through on a streaming platform. How would you read the plain-text reviews from a million different users all at once in a quantum superposition? This bottleneck—the challenge of efficiently accessing the classical world in quantum superposition—is known as the data loading problem. It has arguably been one of the main obstacles to achieving broadly applicable quantum advantage.

Sketching a quantum oracle

In this new work [1], we provide a solution to this seemingly impossible challenge. We develop a framework, called quantum oracle sketching, that enables us to access the classical world in quantum superposition in an optimal way. Importantly, it automatically handles the noise and correlations in the data, and natively supports flexible data structures like vectors and matrices that enable machine learning applications.

The core mechanism relies on processing data as a continuous stream. For each classical data sample we observe, we apply a carefully designed, small quantum rotation to our system. By sequentially accumulating these quantum rotations, we incrementally build up an accurate approximation of the target quantum oracle, which can then be used in any quantum algorithm for data processing. Because every data sample is processed once and immediately discarded, we completely eliminate the massive memory overhead typically required to store the dataset. The fundamental price to pay for assembling quantum queries from classical data lies in the sample complexity: our algorithm consumes a number of samples that scales quadratically with the number of quantum queries we need to make. We show that this rate is optimal and fundamentally arises from the quadratic relationship between quantum amplitudes and classical probabilities governed by the Born rule.

With the data successfully loaded into the quantum computer, the final challenge is to efficiently read out classical results. To address this, we develop an efficient measurement protocol called interferometric classical shadow. Combined with quantum oracle sketching, it allows us to circumvent the data loading and readout bottleneck to construct exponentially compact classical models from massive classical data with quantum technology.

Exponential quantum advantage in machine learning

Using this new approach, we are finally able to find exponential quantum advantage in processing classical data and machine learning. We rigorously prove that a small quantum computer can perform large-scale classification and dimensionality reduction on massive classical data by processing samples on the fly. In contrast, any classical machine achieving the same prediction performance requires exponentially larger size. When the classical machine does not have the required exponentially large memory size, it needs super-polynomially more samples and time relative to our protocol running on a quantum device. Remarkably, this illustrates that quantum technology enables us to construct compact and accurate classical models out of classical data, which is impossible with classical machines alone unless given exponentially larger memory.

The true scale of this exponential memory advantage is staggering. A quantum processor with 300 logical qubits can outperform a classical machine built from every atom in the observable universe. Of course, to actually see such a comical contrast, we would also need universe-scale datasets and processing time.

To contextualize these results in realistic scenarios, consider a large-scale scientific experiment, like a large particle collider. Each experimental run generates a colossal volume of data. With a quantum computer, we can keep squeezing all the data into this tiny quantum chip to perform downstream machine learning tasks such as classification and dimensionality reduction. But if we only have classical machines, we would need to build massive, energy-consuming data centers to store the raw data to match the performance. Without this massive memory overhead, classical machines simply couldn’t extract the same clear signals from a single run, forcing us to repeat the massive, expensive experiment many more times to compensate. To put this into perspective, the Large Hadron Collider (LHC) at CERN generates petabytes (millions of gigabytes) of data per hour, but the data storage bottlenecks force researchers to discard all but a tiny fraction—retaining perhaps only one in a hundred thousand events.

We validated these quantum advantages on real-world datasets, including movie review sentiment analysis and single-cell RNA sequencing. In these public datasets, we demonstrate four to six orders of magnitude (ten thousand to a million times) reduction in memory size with fewer than 60 logical qubits. Given the rapid advancements in high-rate quantum error correction codes and experimental techniques, quantum computers capable of demonstrating such applications are foreseeable in the near future. Crucially, the quantum advantage we propose likely carries a clearer positive impact for society and likely arrives sooner than the applications in cryptanalysis, where the current best estimate requires a thousand logical qubits.

Towards Quantum AI

Our results provide strong evidence that the utility of quantum computers extends far beyond specialized tasks, opening a path for quantum computers to be broadly useful in our everyday life. Rather than fearing that classical AI will “eat quantum computing’s lunch,” we now have rigorous evidence pointing towards a much more exciting prospect: quantum-enhanced AI overpowering classical AI.

Of course, there is still a long way to go towards the dream of quantum intelligence. Our current results establish the provable supremacy of quantum machines in foundational machine learning tasks, such as high-dimensional linear classification and dimensionality reduction. They do not yet imply immediate utility for modern generative AI such as large language models.

That said, our results give me a strong feeling that we are living in an age strikingly reminiscent of the traditional machine learning era—an age dominated by support vector machines and random forests; an age when we relied on rigorous statistical analysis because we lacked the computational resources for large-scale heuristic exploration; an age that ultimately heralded the birth of deep learning and the AI revolution. Today, quantum AI seems to sit at a similar historical position. I cannot wait to see what quantum AI will become once we are capable of unconstrained heuristic exploration on large-scale fault-tolerant quantum computers.

To accelerate this dawn of quantum AI, we invite physicists, computer scientists, developers, and machine learning practitioners to join our efforts and help us push the boundaries of what quantum AI can achieve. To bridge the gap between abstract quantum theory and hands-on machine learning practice, we are open-sourcing our core framework. Our numerical implementation of quantum oracle sketching is built in JAX, natively supporting GPU/TPU acceleration and automatic differentiation to integrate nicely with modern machine learning pipelines. Check out the code, run the simulations, and help us shape the future of quantum AI at github.com/haimengzhao/quantum-oracle-sketching!


References

[1]. Haimeng Zhao, Alexander Zlokapa, Hartmut Neven, Ryan Babbush, John Preskill, Jarrod R. McClean, and Hsin-Yuan Huang. Exponential quantum advantage in processing massive classical data, arXiv:2604.07639, 2026.

May 21, 2026

John PreskillNicole’s guide to handling failure and rejection

It’s happening. 

Your inbox registers an email from the chair of a faculty-hiring committee. With trembling fingers, you click on the message. “We’ve been grateful for the opportunity to learn about your work…The decision was very difficult…many highly qualified candidates…” Months of labor, soul-searching, strain, and anxiety give way to despair. The committee has filled the position, and not with you.

I recently published advice about how to proceed if you receive an offer of a faculty position. But what if you don’t receive an offer—what if hiring committees reject you? Or scholarship committees, grant committees, admissions committees, or potential advisors? What if a journal referee shreds your magnum opus? Or a program committee declines your submission to a conference?

Failure suffuses science as dinner suffuses a nighttime diaper, for two reasons. First, undertaking science is difficult. By “undertaking science,” I mean formulating and proving theorems, cajoling equipment into working, identifying bugs in code, extracting meaning from noisy data, etc. Second, science doesn’t unfold in a vacuum. Human beings do science within a society fraught with opinions, emotions, and limited resources. These challenges lead to the failures and rejections that this article addresses most. (If you’re a member of my group and you’re facing a failure due to the difficulty of undertaking science, come talk with me.) 

What can you do if failure or rejection ails thee? The rest of this article prescribes, in chronological order, steps that will help you recover. You can even turn your lemons into, if not lemonade, then not-entirely-unappetizing lemon meringue pie.

“O what can ail thee, knight-at-arms, alone and palely loitering?”
“The scientific life is rough.”

Immediately after receiving the news:

  • Read the communication once—or, if you must, twice. Don’t linger over the letter for longer than necessary.
  • Put the communication away, so you won’t see it unless you try to. If the news came via email, remove the message from your inbox.
  • Lower your heart rate. Electrified with anger? Expend that energy; go for a walk, for a run, or to the gym, if possible. If you can’t, breathe deeply for several minutes, extending your exhalations.
  • Review records of your successes. I maintain a folder called “Nice messages.” It contains notices of awards I’ve received, kudos on papers I’ve published, messages such as “Thanks for everything this past semester! Your class was probably my favorite one,” and every compliment I’ve received from the taciturn John Preskill via email. Review evidence—remind yourself—that you’re not a failure even though you’re experiencing failure or rejection.

Avoid thinking about the news for a few days. Imagine cutting yourself on a kitchen knife. The wound may sting and bleed initially. But the blood clots and the stinging diminishes if you cover the wound and don’t aggravate it. 

After experiencing a failure or rejection, invite a colleague to lunch, and ask about their recent reading and travels. Dive into a project that will absorb you. Visit a museum, or watch a movie. Remove the letdown from your thoughts.

Consider seeking input from a mentor, especially if you’ve never experienced a failure or rejection of this type before. Mentors have more experience and so can put obstacles in perspective. For example, suppose you’re a student who’s received an upsetting referee report from a journal. The report might look mild to a faculty member, who’s likely received far more, and more-upsetting, reports. What sounds harsh to you might sound quotidian to an advisor, whose lack of distress might reassure you. 

Also, mentors notice upsides that you’ve overlooked. Perhaps the referee trashed your presentation of an idea but tacitly approved of the idea itself. The trashing might have drowned out the approval during your reading. Yet the approval could justify a resubmission to the journal, upending your belief in the cause’s hopelessness. Relatedly, a mentor can help you identify strategies for moving past the failure. Maybe this journal won’t publish your paper but another journal is soliciting contributions to a relevant special collection.

As a PhD student, I applied for an internship at a company developing a quantum computer. I broke an obligation and flew out of state to interview for the position. No offer materialized. To process the outcome, I spoke with a more-advanced researcher I trusted. He not only offered a mature perspective, but also had access to inside information. The team liked me, he reported, but my interests didn’t overlap with theirs enough. In a sense, I’d grown too independent. As independence marks maturity in PhD students, the rejection came to double as a compliment. Today, I remain on friendly terms with multiple people from that team, and a student of mine just landed an internship at a quantum-computing company.

Someone more experienced than you should have your back. (For the story behind this photo, see this blog post.)

Chart a path past the failure or rejection. Cue the lemon meringue pie. Did a hiring committee reject an application of yours? Email the committee’s chair, thanking them for considering your application. Express your hope of improving your materials so that you can try again the following year. Ask if the chair would provide feedback via phone or Zoom.1 

Pursue the path you’ve charted. To ease the burden, consider undertaking lighter tasks before more-arduous ones. I recently received a laundry list of requests about a manuscript from an editor—and when I say laundry, I mean the equivalent of hauling a multiple-pound bag to a stream, scrubbing everything by hand while the wind attempts to blow cleaned items into the mud, and then hauling everything back. I began with the tasks that required little effort. Dispatching them, I crossed about half the requests off my to-do list. The list looked more manageable as a result; I felt better-equipped to handle the trickiest work.

Celebrate your triumphs. Don’t let failures and rejections consume your attention; carve out time for your successes. If you recognize them, you’ll remember them the next time you face failure; they’ll cushion you when you fall.

I recently failed at a task over a hundred times (yes, I counted). After 28 months of trying, I succeeded. I celebrated by treating myself to lunch at the National Gallery of Art.2 Why not gather ye rosebuds while ye may? This success afforded me the opportunity to fail at a new task.

The National Gallery of Art sold a fairly steampunk book near its Garden Café recently.

Always remember, what matters in the long run is not any one failure or rejection, but your persistence despite failures and rejections. I progressed to the Rhodes Scholarship competition’s final round two years in a row. Each year, a committee interviewed and rejected me. I’d poured time, sweat, and blood into my applications and interview preparations. I felt like I had no more blood to squeeze into the applications I then had to write for graduate programs because of those rejections. Those rejections, though, led me to become a student of John Preskill’s. Over a decade later, I’m not kicking myself.

Result of two failures.

Share about your favorite failures in the comments section below!

1Not via email, which doesn’t offer the same freedom.

2Its Garden Café had enticed me for years, but I can rarely bring myself to pay for food when I can prepare it myself.

May 16, 2026

Jacques Distler Code

I’ve been playing around with Claude Code (Claude Opus 4.7 (1M context)) because, well, who hasn’t?

I am, so far, only moderately impressed by its abilities in physics. The most impressive bit so far was when, in response to a question about nilpotent orbits, it responded

I don’t know. I could waste your time by guessing, but …

I had never had an LLM tell me that it doesn’t know something, much less that it didn’t want to waste my time by making stuff up. So this was positively shocking to read.

Claude Code is, however, unfathomably good at generating code, so I set it the task of modernizing Instiki and Heterotic Beast, my forum software. Both are Rails applications and both have extensive test suites. So they use a software framework Claude is familiar with and have an objective standard for whether the changes made are correct.

When there is no test, however, things can go wildly off the rails (pun intended). For instance, consider the following snippet of Ruby code


def foo(text)
  ...
  con = text
  ...
  (now mutate con)
  ...
  con
end

If text is a frozen string, this will generate an error, as you can’t mutate a frozen string. Obviously, what you should do is write


def foo(text)
  ...
  con = text.dup
  ...
  con
end

which copies the caller’s string to a new unfrozen string which you can mutate to your heart’s content.

What did Claude do?


def foo(text)
  ...
  con = text.encode
  ...
  con
end

which also produces a new unfrozen string, transcoded from the caller’s encoding to Encoding.default_internal (which turns out to be nil). This is both (a) nondeterministic and (b) blows up spectacularly when text contains astral plane characters, like “𝔸”. I had to tell Claude not to do that, and to write some tests to check that astral plane characters are handled correctly.

Still …

I would set Claude the task of rewriting this blogging software, but alas I don’t have a test suite to compare with.

What I really should do, though, is find some physics I would trust it to work on.

May 09, 2026

Tim GowersA recent experience with ChatGPT 5.5 Pro

We are all having to keep revising upwards our assessments of the mathematical capabilities of large language models. I have just made a fairly large revision as a result of ChatGPT 5.5 Pro, to which I am fortunate to have been given access, producing a piece of PhD-level research in an hour or so, with no serious mathematical input from me.

The background is that, as has been widely reported, LLMs are now capable of solving research-level problems, and have managed to solve several of the Erdős problems listed on Thomas Bloom’s wonderful website. Initially it was possible to laugh this off: many of the “solutions” consisted in the LLM noticing that the problem had an answer sitting there in the literature already, or could be very easily deduced from known results. But little by little the laughter has become quieter. The message I am getting from what other mathematicians more involved in this enterprise have been saying is that LLMs have got to the point where if a problem has an easy argument that for one reason or another human mathematicians have missed (that reason sometimes, but not always, being that the problem has not received all that much attention), then there is a good chance that the LLMs will spot it. Conversely, for problems where one’s initial reaction is to be impressed that an LLM has come up with a clever argument, it often turns out on closer inspection that there are precedents for those arguments, so it is still just about possible to comfort oneself that LLMs are merely putting together existing knowledge rather than having truly original ideas. How much of a comfort that is I will not discuss here, other than to note that quite a lot of perfectly good human mathematics consists in putting together existing knowledge and proof techniques.

I decided to try something a little bit different. At least in combinatorics, there are quite a lot of papers that investigate some relatively new combinatorial parameter that leads naturally to several questions. Because of the sheer number of questions one can ask, the authors of such papers will not necessarily have the time to spend a week or two thinking about each one, so there is a decent probability that at least some of them will not be all that hard. This makes such papers very valuable as sources of problems for mathematicians who are doing research for the first time and who will be hugely encouraged by solving a problem that was officially open. Or rather, it used to make them valuable in that way, but it looks as though the bar has just been raised. It is no longer enough that somebody asks a problem: it needs to be hard enough for an LLM not to be able to solve it.

In any case, a little over a week ago I decided to see how ChatGPT 5.5 Pro would fare with a selection of problems asked by Mel Nathanson in a paper entitled Diversity, Equity and Inclusion for Problems in Additive Number Theory. Nathanson has a remarkable record of being interested in problems and theorems that have later become extremely fashionable, which has led him to write a series of extremely well timed and therefore highly influential textbooks. In this paper, he argues for the interest of several other problems, some of which I will now briefly describe.

If A is a set of integers, then its sumset A+A is defined to be \{a+b:a,b\in A\}. For a positive integer h, the hfold sumset, denoted hA, is defined to be \{a_1+\dots+a_h: a_1,\dots,a_h\in A\}. Nathanson is interested in the possible sizes of hA given the size of A. To that end one can define a set \mathcal R(h,k) to be the set of all t such that there exists a set A with |A|=k and |hA|=t.

An obvious first question to ask is simply “What is \mathcal R(h,k)?” When h=2, the answer is the set of all integers between 2k-1 and \binom{k+1}2. It is an easy exercise to show that if |A|=k, then 2k-1\leq|A+A|\leq\binom{k+1}2, so this result is saying that all sizes in between can be realized. However, it is not true in general that hA can take every size between its minimum and maximum possibilities, and we do not currently have a complete description of \mathcal R(h,k).

Another natural question one can ask, and this is where ChatGPT came in, is how large a diameter you need if you want a set A with A and hA having prescribed sizes. (Of course, the size of hA must belong to \mathcal R(h,k).) Nathanson showed that for every t\in[2k-1,\binom{k+1}2] there is a subset A of \{0,1,2,\dots,2^k-1\} with |A|=k and |A+A|=t, and asked whether the bound 2^k-1 could be improved. ChatGPT 5.5 Pro thought for 17 minutes and 5 seconds before providing a construction that yielded a quadratic upper bound, which is clearly best possible. It wrote up its argument in a slightly rambling LLM-ish style, so I asked if it could write the argument up as a LaTeX file in the style of a typical mathematical preprint. After two minutes and 23 seconds it gave me that, after which I spent some time convincing myself that the argument was correct.

The basic idea behind both Nathanson’s argument and ChatGPT’s was that in order to obtain a set of a given size with a sumset of a given size, it is useful to build it out of a Sidon set, which means a set with sumset of maximal size (that is not quite the usual definition but it is the simplest to use in this discussion), and an arithmetic progression. Also, for a bit of fine tuning one can take an additional point near the arithmetic progression. Then if one plays around with the various parameters, one finds that one can obtain sets of all the sizes one wants. Nathanson doesn’t express his argument this way (it is Theorem 5 of this paper), instead giving an inductive argument, but I think, without having checked too carefully, that if one unravels his argument, one finds that effectively that is what he ends up with, and the Sidon set in question consists of powers of 2. ChatGPT obtained its improvement by simply using a more efficient Sidon set — it is well known that one can find Sidon sets of quadratic diameter. (One might ask why Nathanson didn’t do that in the first place: I think it is because the obvious idea of using a more efficient Sidon set becomes obvious only after one has redescribed his inductive construction. Is that what ChatGPT did? It is very hard to say.)

Next, I asked ChatGPT to see whether it could do the same for a closely related question, where instead of looking at the size of the sumset, one looks at the size of the restricted sumset, which is defined to be \{a+b:a,b\in A, a\ne b\}. Unsurprisingly, it was able to do that with no trouble at all. I got it to write both results up in a single note, to avoid a certain amount of duplication. If you are curious, you can see the note here.

I then asked what it could do for general h. I was much less optimistic that it would manage to do anything interesting, because the proof for h=2 makes fundamental use of the fact (due to Erdős and Szemerédi) that we know exactly which sizes we need to create. If we don’t know what the set \mathcal R(h,k) is, then it seems that we are forced to start with a hypothetical set A with |A|=k and |hA|=t and build out of it a set of small diameter with the same property. As it happens, I still don’t know how to get round that difficulty (I’m mentioning that just to demonstrate that my mathematical input was zero, and I didn’t even do anything clever with the prompts), but Nathanson mentioned in his paper a remarkable paper of Isaac Rajagopal, a student at MIT, who must have got round the difficulty somehow, because he had managed to prove an exponential dependence of \mathcal R(h,k) on k for each fixed h.

I’ll leave the previous paragraph there, but Isaac has subsequently explained to me that that isn’t really the difficulty. His argument gives a complete description of \mathcal R(h,k) when k is sufficiently large, and if one wants to prove a polynomial dependence for fixed h, then assuming that k is sufficiently large is clearly permitted. The real difficulty is that constructing the sets with given sumset sizes was significantly more complicated, and necessarily so because the degree of the polynomial grows with h, and one therefore needs more and more parameters to define the sets.

In any case, the task faced by ChatGPT was not to solve the problem from scratch, but to see whether it was possible to tighten up Isaac Rajagopal’s argument. Here’s what happened.

  1. After 16 minutes and 41 seconds, it came back with an argument that claimed to have improved the upper bound from exponential in k to exponential in k^\alpha for any \alpha>1/2.
  2. I asked it to write that in preprint form too, which took it a further 47 minutes and 39 seconds.
  3. That preprint would have been hard for me to read, as that would have meant carefully reading Rajagopal’s paper first, but I sent it to Nathanson, who forwarded it to Rajagopal, who said he thought it looked correct.
  4. Both ChatGPT and Rajagopal speculated a little on what might need to be done to push things further and get a polynomial bound, so I got greedy and asked ChatGPT to give that a go.
  5. After 13 minutes and 33 seconds it told me it felt optimistic about the existence of such an argument but there were a couple of technical statements that needed checking.
  6. I asked it to check them.
  7. After 9 minutes and 12 seconds it got back to me with the check having been done, so I asked for this too to be written in preprint form.
  8. After 31 minutes and 40 seconds the “preprint” was ready. Here it is.
  9. Isaac Rajagopal looked at it and declared it to be almost certainly correct. It was clear that he meant this not just at a line-by-line level but at the level of ideas.

Isaac made some very interesting remarks about the nature of what the additional ideas were that ChatGPT contributed. Since, as I have already said, my mathematical input was zero, I invited him to write a guest section to this post. Just before we get to that, I want to raise a question (that will undoubtedly have been raised by others as well), which is simple: what should we do with this kind of content? Had the result been produced by a human mathematician, it would definitely have been publishable, so I think it would be wrong to describe it as AI slop. On the other hand, it seems pointless even to think about putting it in a journal, since it can be made freely available, and nobody needs “credit” for it (except that Isaac deserves plenty of credit for creating the framework on which ChatGPT could build). I understand that arXiv has a policy against accepting AI-written content, which makes good sense to me. So maybe there should be a different repository where AI-produced results can live. But various decisions would need to be made about how it was organized. I myself think that one would probably want to have some kind of moderation process, so that results would be included only if a human mathematician was prepared to certify that they were correct — or, better still, that they had been formalized by a proof assistant — and perhaps also that they answered a question that had been asked in a human-written paper. On the other hand, I wouldn’t want a moderation process that created vast amounts of work (unless the work was itself done by AI, but there are obvious dangers in going down that route). Anyway, until these questions are answered, this result is available from the link above, and perhaps, now that LLMs are so good at literature search, that will be enough to make it findable by anyone who wants to know whether Nathanson’s problem has been solved.

Isaac’s evaluation of what ChatGPT achieved

With just a few prompts, ChatGPT was able to improve the upper bound on N(h,k) (which I will define very soon) from exponential in k to polynomial in k. While its first improvement of the bound, from exponential in k to exponential in k^{\frac{1}{2} + \varepsilon}, was a routine modification of my work, the improvement to polynomial in k is quite impressive. To do this, ChatGPT came up with an idea which is original and clever. It is the sort of idea I would be very proud to come up with after a week or two of pondering, and it took ChatGPT less than an hour to find and prove, using similar methods to those in my own proof. My goal is to explain that idea, in a manner that will be digestible to my friends who are computer science majors as well as my math major friends.

The problem of bounding N(h,k) is closely related to a problem I worked on at the Duluth REU (Research Experience for Undergrads) program, of determining \mathcal{R}(h,k). In particular, \mathcal{R}(h,k) is the set of possible h-fold sumset sizes |hA|, where A can be chosen to be any set of k integers. N(h,k) is the minimal N such that we can achieve all of the values of \mathcal{R}(h,k) using k-element sets A \subset \{0,1,2,\ldots,N\}. I spent last summer explicitly characterizing the set \mathcal{R}(h,k) for large k, by constructing sets A such that |hA| achieves all sizes which I could not rule out as impossible. So, N(h,k) can be upper-bounded by optimizing my constructions.

I constructed these sets A by combining smaller component sets which are simpler to analyze. Some of these components are the geometric series

\displaystyle S = \{0,1,m,m^2,\ldots,m^{\ell-2}\} \quad \hbox{and} \quad T = \{1,m,m^2,\ldots,m^{\ell-1}\} \qquad (1)

for various values of 2 \leq m \leq h and 2 \leq \ell \leq k. Unfortunately, the elements of S and T are exponentially large in terms of k. So, I asked ChatGPT (through Tim) whether there exist sets of \ell elements which have similar sumset sizes to these geometric series, but contain only numbers of polynomial size in \ell: I had no idea if this was possible, or how to begin constructing such sets. ChatGPT came back with an answer, constructing sets G and H which behave like “half a geometric series squeezed into a polynomial interval,” which is counterintuitive. Before I discuss the construction of G and H, I will explain the important properties of the sumset sizes of S and T which they recreate.

For h > 0, a set A is called a B_h set if the only solutions to

\displaystyle x_1+\cdots+x_h = y_1+\cdots+y_h

with x_i,y_i in A are the “trivial” solutions, by which I mean that one side of the equation is a reordering of the other side. If A is a B_h set of size \ell, then elements of hA correspond exactly to choices of h elements of A, with repetition allowed. Using “stars and bars,” one can see that |hA| = \binom{h+\ell - 1}{h} and this is the maximum possible value of |hA| among sets of size \ell. So, another definition is that A is a B_h set if |hA| = \binom{h+|A| - 1}{h}. Sidon sets, which Tim discussed, are exactly B_2 sets.

To make things more concrete, let us assume that m = 4 in (1). Then, S is a B_3 set, but it is not a B_4 set because of the relations

\displaystyle 4^{a} + 4^a + 4^a + 4^a = 4^{a+1} + 0 + 0 + 0 \qquad (2)

for any choice of a in \{0,1,2,\ldots, \ell-3\}. In particular, \binom{\ell+3}{4} - |4S| = \ell-2, as these \ell-2 relations are the only ones preventing S from being a B_4 set. T lacks the relations in (2) because 0 is not in T. So, T is a B_4 set, but it is not a B_5 set because of the relations

\displaystyle 4^{a} + 4^a + 4^a + 4^a + 4^{b+1} = 4^{a+1} + 4^b + 4^b + 4^b + 4^b \qquad (3)

for any choices of a \neq b in \{0,1,2,\ldots, \ell-2\}. This gives \binom{\ell-1}{2} relations, and one can check that \binom{\ell+4}{5} - |5T| = \binom{\ell-1}{2}. To summarize, we have seen that

(a) S is a B_{m-1} set.

(b) \binom{m+\ell-1}{m} - |mS| = \ell -2 is a linear function of \ell.

(c) T is a B_{m} set.

(d) \binom{m+\ell}{m+1} - |(m+1)T| = \binom{\ell-1}{2} is a quadratic function of \ell.

    ChatGPT was able to find sets G and H of \ell elements which satisfy (a)-(d), but whose elements all have polynomial size in \ell. The construction of G and H uses h^2-dissociated sets, which are sets A where the only solutions to

    \displaystyle x_1+\cdots+x_s = y_1+\cdots+y_{s'} \qquad (4)

    with s,s' \leq h^2 and x_i,y_i in A are the “trivial” solutions, i.e. s = s' and one side of the equation is a reordering of the other side. For r > 0, it is possible to construct an h^2-dissociated set U = \{u_1,\ldots,u_r\} \subseteq \{0,1,2,\ldots,N\}, where N is approximately r^{h^2}, and in particular polynomial in r. Constructions of such a U using finite fields date back to Singer (1938) and Bose–Chowla (1963) and are described in Appendix 1. Define

    \displaystyle G = \{0, u_1,u_2,\ldots,u_r,mu_1,mu_2,\ldots, mu_r\}

    and

    \displaystyle H= \{u_1,u_2,\ldots,u_r,mu_1,mu_2,\ldots, mu_r\}. \qquad (5)

    In hindsight, I have good intuition for the construction of G and H. All of the relations in (2) and (3) are formed by combining one or two relations of the form 4x = y. There are approximately \ell relations of the form mx = y in S and T, and approximately \ell/2 such relations in G and H. There are few other low-order relations in S and T, and similarly in G and H because U is h^2-dissociated. So, G and H manage to contain half as many mx = y-relations as their geometric series counterparts, while also containing few low-order relations.

    We now see why (a)-(d) hold with S and T replaced by G and H, respectively. For concreteness, we assume that m = 4 and h>4, so U contains no nontrivial relations as in (4) with s,s' \leq 25 \leq h^2. Then, G is a B_3 set, but it is not a B_4 set because of the relations

    \displaystyle u_i + u_i + u_i + u_i = 4u_i + 0 + 0 + 0

    for any choice of i in \{1,2,\ldots, r\}. If we let \ell = |G| = 2r+1, we can check that \binom{\ell + 3}{4} - |4G| = r = \frac{\ell-1}{2} is linear in \ell. In particular, (a) and (b) hold with S replaced by G, and the linear function \ell-2 replaced by \frac{\ell-1}{2}. We can also see that H is a B_4 set, but it is not a B_5 set because of the relations

    \displaystyle u_i + u_i + u_i + u_i + 4u_j = 4u_i + u_j + u_j +u_j + u_j

    for any i\neq j in \{1,2,\ldots, r\}. If we let \ell = |H| = 2r, we can check that \binom{\ell + 4}{5} - |5H| = \binom{r}{2} = \binom{\ell/2}{2} is quadratic in \ell. In a similar manner, (c) and (d) hold with T replaced by H, and the quadratic function \binom{\ell-1}{2} replaced by \binom{\ell/2}{2}.

    Even though I can motivate it in retrospect, ChatGPT’s idea to use h^2-dissociated sets to control relations of order at most h feels quite ingenious. As far as I can tell, this idea is completely original.

    ChatGPT’s proof that its construction produces the desired values of |hA| is very similar to my proof that the sets A which I construct achieve all possible values of |hA|, after replacing S and T by G and H, respectively. Properties (a)-(d) capture many of the important properties of S and T (or G and H) which are used in this proof. The final constructions involve combining the sets G and H (or S and T in my paper) for each value of m between 2 and h with another set which is the union of an arithmetic progression and a point. Intuitively, G and H (or S and T) have large sumsets, while arithmetic progressions have small sumsets, so it is plausible that one could get sets which achieve all the medium-sized sumsets by combining them. However, the proof of this is quite involved, and it occupies Section 4 of my paper and the entirety of the ChatGPT preprint. In Appendix 2, I work out the details of the ChatGPT construction to show that for k sufficiently large,

    \displaystyle N(h,k) \leq O\left(k^{10h^3}\right).

    For comparison, it is easy to see that N(h,k) is at least on the order of k^{h}, and it is unknown what the real value is. In Appendix 3, I give details of the correspondence between my paper and the ChatGPT preprint, which will be helpful for those who want to read either.

    Finally, I want to express my deep gratitude to Tim for allowing me to contribute to this blog. I am still stunned by the coincidence that the problem he chose to put into ChatGPT 5.5 Pro led him to my paper on the arXiv.

    Tim on what this means for mathematical research

    I would judge the level of the result that ChatGPT found in under two hours to be that of a perfectly reasonable chapter in a combinatorics PhD. It wouldn’t be considered an amazing result, since it leant very heavily on Isaac’s ideas, but it was definitely a non-trivial extension of those ideas, and for a PhD student to find that extension it would be necessary to invest quite a bit of time digesting Isaac’s paper, looking for places where it might not be optimal, familiarizing oneself with various algebraic techniques that he used, and so on.

    It seems to me that training beginning PhD students to do research, which has always been hard (unless one is lucky enough, as I have often been, to have a student who just seems to get it and therefore doesn’t need in any sense to be trained), has just got harder, since one obvious way to help somebody get started is to give them a problem that looks as though it might be a relatively gentle one. If LLMs are at the point where they can solve “gentle problems”, then that is no longer an option. The lower bound for contributing to mathematics will now be to prove something that LLMs can’t prove, rather than simply to prove something that nobody has proved up to now and that at least somebody finds interesting.

    I would qualify that statement in two ways though. First, there is the obvious point that a beginning PhD student has the option of using LLMs. So the task is potentially easier than proving something that LLMs can’t prove: it is proving something in collaboration with LLMs that LLMs cannot manage on their own. I have done quite a lot of such collaboration recently and found that LLMs have made useful contributions without (yet) having game-changing ideas.

    A second point is that I don’t know how much of what I have said generalizes to other areas of mathematics. Combinatorics tends to be quite focused on problems: you start with a question and you reason back from the question or if you reason forwards you do so very much with the question in mind. In other areas there can be much more of an emphasis on forwards reasoning: you start with a circle of ideas and see where it leads. To do it successfully, you need to have some way of discriminating between interesting observations and uninteresting ones, and it isn’t obvious to me what LLMs would be like at that.

    Of course, everything I am saying concerns LLMs as they are right now. But they are developing so fast that it seems almost certain that my comments will go out of date in a matter of months. It is also almost certain that these developments will have a profoundly disruptive effect on how we go about mathematical research, and especially on how we introduce newcomers to it. Somebody starting a PhD next academic year will be finishing it in 2029 at the earliest, and my guess is that by then what it means to undertake research in mathematics will have changed out of all recognition.

    I sometimes get emails from people who are interested in doing mathematical research but are not sure whether that makes sense any more as an aspiration. I have a view on that question, but it may very well change in response to further developments. That view is that there is still a great deal of value in struggling with a mathematics problem, but that the era where you could enjoy the thrill of having your name forever associated with a particular theorem or definition may well be close to its end. So if your aim in doing mathematics is to achieve some kind of immortality, so to speak, then you should understand that that won’t necessarily be possible for much longer — not just for you, but for anybody. Here’s a thought experiment: suppose that a mathematician solved a major problem by having a long exchange with an LLM in which the mathematician played a useful guiding role but the LLM did all the technical work and had the main ideas. Would we regard that as a major achievement of the mathematician? I don’t think we would.

    So what is the point of struggling with a difficult mathematics problem? One answer is that it can be very satisfying to solve a problem even if the answer is already known, but I don’t think that is a sufficient reason to spend several years of your life on this peculiar activity. A better answer is that by solving hard problems you get an insight into the problem-solving process itself, at least in your area of expertise, in a way that you simply don’t if all you do is read other people’s solutions. One consequence of this is that people who have themselves solved difficult problems are likely to be significantly better at using solving problems with the help of AI, just as very good coders are better at vibe coding than not such good coders, or people who have a solid grasp of how to do basic arithmetic are likely to be more skilled at using calculators (and especially at noticing when an answer feels off). Mathematics is a highly transferable skill, and that applies to research-level mathematics as well. By doing research in mathematics, you may not get the same rewards as your equivalents a generation ago, but there is a good chance that you will be equipping yourself very well for the world we are about to experience.

    Appendix 1 (Isaac)

    We will construct an h-dissociated set U = \{u_1,\ldots,u_r\} \subseteq \{0,1,2,\ldots,N\}, where N is approximately r^{h}. This construction is a very minor modification of Bose–Chowla (1963)’s construction of a B_h set, which I learned about from this paper. For whatever reason, the GPT preprint (Lemma 3.1) uses a different, less efficient construction using moment curves.

    Let p > r be a prime, let N = p^{h+1}-2, let K be the finite field with p^{h+1} elements and fix a generator \theta of K^\times, so that K^\times is equal to \{\theta^0,\theta^1,\ldots, \theta^N\}. Define a set of p elements

    \displaystyle U = \{a \in \{0,1,2,\ldots,N\}: \theta^a - \theta \in \mathbb{F}_p\}.

    Then, each element a \in U corresponds to a unique value of \tilde{a} \in \mathbb{F}_p, by taking \tilde{a} = \theta^a - \theta. Now an additive relation of the form in (4) with s,s' \leq h can be reframed by taking powers of \theta as

    \displaystyle (\theta + \tilde{x_1})(\theta + \tilde{x_2})\cdots (\theta + \tilde{x_s}) = (\theta + \tilde{y_1})(\theta + \tilde{y_2})\cdots (\theta + \tilde{y_{s'}}). \qquad (6)

    As K is a degree-h+1 extension of \mathbb{F}\sb{p} and \theta is a generator of K as an \mathbb{F}\sb{p}-extension, this means that \theta does not satisfy any nonzero polynomials in \mathbb{F}\sb{p}[x] of degree \leq h. So, both sides of (6) are identical as polynomials in \mathbb{F}_{p}[\theta] and thus the additive relation in (4) is trivial. So, U is h-dissociated, and of course one can prune a few elements to reduce U to size r.

    Appendix 2 (Isaac)

    Fix constants \alpha,\beta,\gamma such that 0.5 < \beta\gamma < \beta < \alpha < 1 (in my paper I arbitrarily chose (\alpha,\beta,\gamma) = (0.9,0.8,0.7)). Let the two sets in (5) be called G_{m,r} and H_{m,r}. Let [a,b] denote the set of integers x satisfying a \leq x \leq b. Similarly to my paper, the constructions of A such that hA achieves the desired sizes will combine sets of the following four types:

    • B_{j,b} := [0,b-2] \cup \{b-2+j\} with choices of b \in [3, k-k^\gamma] and j \in [1,hb].
    • G_{m,r_m} for each value of m \in [3, h], with choices of r_m \in [0, (k-b)^\alpha].
    • H_{m,u_m} for each value of m \in [2,h-1], with choices of u_m \in [0, (k-b)^\beta].
    • A B_h set of the correct size so that |A| = k.

    One reason that this construction needs to be complicated is that we need to create at least \Omega(k^h) many sets. To do this, we vary 2h-4 parameters r_m and u_m in the domain [0,k^\alpha] and 2 parameters b and j in the domain [1,hk]. We can choose \alpha to be slightly bigger than 1/2, and then the above construction gives us O(k^{\alpha(2h-4)+ 2})=O(k^{h + \delta}) different sets where \delta >0 can be made arbitrarily small. So, if we were to remove any of the above parameters from the construction, and not change the others, this construction would no longer create \Omega(k^h) many sets. In comparison, Nathanson’s construction when h=2 only needs to create \Omega(k^2) sets. He does this by combining a Sidon set, an arithmetic progression, and one extra value, and varying the size of the arithmetic progression and the extra value in ranges of size O(k).

    We want to combine q = 2h-2 sets A_1,\ldots,A_q, which are given by B_{j,b}, G_{m,r_m} for the h-2 values of m \in [3,h], H_{m,u_m} for the h-2 values of m \in [2,h-1], and a B_h set. By Appendix 1, for all r \leq k, there exists a h^2-dissociated set {u_{1},\ldots,u_{r}} of diameter M \leq r^{2h^2} \leq k^{2h^2}. By the constructions of G_{m,r_m} and H_{m,u_m}, we can take each A_i \subseteq [0,M], where M \leq hk^{2h^2}. Let \mathbb{Z}^{2q} have basis vectors e_1,\ldots,e_{2q}. To combine A_1,\ldots,A_q, we can define A \subseteq \mathbb{Z}^{2q} as

    \displaystyle A = \bigcup_{i=1}^q (A_i e_i + e_{q+i}) \subseteq \{0,1,2,\ldots,M\}^{2q} \subseteq \mathbb{Z}^{2q}.

    Similarly to my Lemma 4.9, this construction ensures that the generating function product \mathcal{F}_{A}(z) = \prod_{i=1}^q \mathcal{F}_{A_i}(z) holds, which is the identity that both my paper and the GPT preprint use (see either paper for a definition of these generating functions). By (the standard) Lemma 2.3 of the GPT preprint, A is Freiman-isomorphic of order h to a subset of [0,2qM(2hM)^{2q-1}]. Therefore, for k sufficiently large (the whole construction relies on this for the same reasons as in my paper),

    \displaystyle N(h,k) \leq 2qM(2hM)^{2q-1} \leq 2\left(2h^2k^{2h^2}\right)^{2(2h-2)} \leq k^{10 h^3}.

    Appendix 3 (Isaac)

    In Section 4.2 of my paper, I use a different, simpler construction to construct sets A achieving the values in \mathcal{R}(h,k) which have |hA| < \varepsilon k^h, for some small \varepsilon. These sets A are subsets of {0,1,2,\ldots,k^h}, meaning that all elements have polynomial size in k. This is observed in Section 5 of the GPT preprint.

    Section 4.3 of my paper carries out the construction which combines many components including S and T. This corresponds to Sections 2, 3, 4, and 6 of the GPT preprint. This section has a lot of moving parts; I give an outline in Section 4.3.1.

    In Section 4.3.2, I describe how the different components will be combined, using a construction which I call the disjoint union, and introduce generating functions \mathcal{F}_A(z) as a bookkeeping tool to keep track of the sumset sizes of a set A. This corresponds to Section 2 and Section 4 of the GPT preprint.

    In Section 4.3.3, I compute the generating function of each of the component sets, including \mathcal{F}_S(z) (Lemma 4.15) and \mathcal{F}_T(z) (Lemma 4.17). This corresponds to Section 3 and Section 6.1 of the GPT preprint. In particular, \mathcal{F}_{G}(z) is computed in Lemma 3.3 and \mathcal{F}_{H}(z) is computed in Lemma 3.4. Once these generating functions have been computed, the remainder of the proof is almost identical in my paper and in the GPT preprint.

    In Section 4.3.4, I put all the pieces together to show that as we range over the sets A which I have constructed, the values of |hA| will assume all of the elements of {\lceil\varepsilon k^h\rceil, \lceil\varepsilon k^h\rceil+1,\ldots ,\binom{h+k-1}{h} }. The key idea is to show that the set of all values of |hA| forms an interval, and contains numbers both smaller than \varepsilon k^h and equal to \binom{h+k-1}{h}.

May 02, 2026

n-Category Café Quantum Mechanics of the Inverse Cube Force Law

In the last episode of my column in Notices of the American Mathematical Society, we looked at a particle moving in an attractive central force whose strength is proportional to the inverse cube of the distance from the origin. Among other things, we saw that a particle moving in such a force can spiral in to the origin in a finite time. But that was classical mechanics. What about quantum mechanics?

Here things get more tricky. The uncertainty principle tends to prevent the particle from falling in to the origin. But when the attractive force is strong enough, the particle can still fall in. We can make up a theory where the particle shoots back out, but there are choices involved: we need to say how the particle changes phase when shoots back out. So there is not just a single theory, but many!

Why does the particle come back out? There are theories where it does not. In these theories, at least those studied so far, time evolution is nonunitary: that is, the probability of finding the particle somewhere or other does not stay equal to 11, because the particle simply disappears when it hits the origin. Here we focus on theories where time evolution is unitary and the particle comes back out. Many people have written about these, running into ‘paradoxes’ when they weren’t careful enough. Only rather recently have things been straightened out.

Let us dig into the details. In quantum mechanics, the Hilbert space of states of a particle in 3\mathbb{R}^3 is L 2( 3)L^2(\mathbb{R}^3). In a central force whose strength is proportional to 1/r 31/r^3, such a particle has a Hamiltonian of this form:

H= 2+cr 2 H = -\nabla^2 + c r^{-2}

The first term describes the particle’s kinetic energy, while the second describes its potential energy: remember, taking the gradient of an inverse square potential gives an inverse cube force. I have set some constants to 11 to remove irrelevant clutter, but we need the constant cc to say how strong the force is. When c<0c \, &lt; \, 0, the force is attractive.

In this game, analysis is paramount. We should interpret HH as a densely defined linear operator on L 2( 3)L^2(\mathbb{R}^3). For this, we choose a dense linear subspace DL 2( 3)D \subset L^2(\mathbb{R}^3) and treat HH as a linear map from DD to L 2( 3)L^2(\mathbb{R}^3). Different choices of DD correspond to different physical assumptions: for example, assumptions about what happens when the particle falls into the origin.

To get unitary time evolution in quantum mechanics, we need the Hamiltonian to be self-adjoint. But adjoints of densely defined operators are tricky. Let us briefly recall how they work. Given a Hilbert space \mathcal{H} and a linear operator AA from a dense linear subspace D(A)D(A) \subseteq \mathcal{H} to \mathcal{H}, we define D(A *)D(A^*) to be the set of all ψ\psi \in \mathcal{H} for which there exist ψ\psi' \in \mathcal{H} such that

ψ,ϕ=ψ,Aϕ for all ϕD(A). \langle \psi' , \phi \rangle = \langle \psi, A \phi \rangle \; \text{ for all } \; \phi \in D(A).

If such a vector ψ\psi' exists, it is unique, and it depends linearly on ψ\psi. Thus, for ψD(A*)\psi \in D(A\ast) we define A*ψA\ast \psi to be the vector ψ\psi' with the above property. The adjoint of AA is then the linear operator A*:D(A*) A\ast \colon D(A\ast) \to \mathcal{H}. We say AA is self-adjoint if A=A*A = A\ast. We say that AA is essentially self-adjoint if it has a unique extension to a self-adjoint operator. If it does, this extension must be A*A\ast.

All this raises the question of whether the Hamiltonian HH for the inverse cube force law can be made self-adjoint with a suitable choice of domain. It turns out we can always do it, but sometimes in more than one way. There are three regimes:

  • c34c \ge \tfrac{3}{4}. In this case we can start with the domain C 0 ( 3{0})C_0^\infty(\mathbb{R}^3 - \{0\}) consisting of smooth functions that are compactly supported on 3\mathbb{R}^3 minus the origin. The operator HH is unambiguously defined on this domain, and it is essentially self-adjoint.

  • 14c<34-\tfrac{1}{4} \le c \, &lt; \, \tfrac{3}{4}. In this case HH is still well-defined on the domain C 0 ( 3{0})C_0^\infty(\mathbb{R}^3 - \{0\}), but it is not essentially self-adjoint. In fact, it admits more than one self-adjoint extension! However, HH is bounded below: there is a constant E 0E_0 such that ψ,HψE 0ψ,ψ \langle \psi, H \psi \rangle \ge E_0 \langle \psi, \psi \rangle for all ψC 0 ( 3{0})\psi \in C_0^\infty(\mathbb{R}^3 - \{0\}). Physically, this means that the particle’s energy is bounded below by E 0E_0. Mathematically, this implies that HH has a canonical choice of self-adjoint extension called the ‘Friedrichs extension’, with the smallest possible domain. But there is another canonical choice, the ‘Krein extension’, with the largest possible domain.

  • c<14c \, &lt; \, -\tfrac{1}{4}. In this case HH is well-defined on the domain C 0 ( 3{0})C_0^\infty(\mathbb{R}^3 - \{0\}), and it has more than one self-adjoint extension, but it is not bounded below.

These strange results demand explanation. For example, what is special about c=14c =-\tfrac{1}{4}? In classical mechanics, the energy of a particle in the inverse cube force ceases to be bounded below as soon as c<0c \, &lt; \,0. Quantum mechanics is different. To get a lot of negative potential energy, the particle’s wavefunction must be peaked near the origin, but that gives it kinetic energy. The tradeoff is captured by Hardy’s inequality. This says that for any ψC 0 ( 3)\psi\in C_0^\infty(\mathbb{R}^3) we have

ψ,( 214r 2)ψ0. \langle \psi, (-\nabla^2 - \tfrac{1}{4} r^{-2}) \psi \rangle \ge 0 .

This is why HH is bounded below when c14c \ge -\tfrac{1}{4}.

On the other hand, the constant 14\tfrac{1}{4} in Hardy’s inequality cannot be improved, so if c<14c \, &lt; \, \tfrac{1}{4} we can find ψ\psi with ψ,Hψ<0 \langle \psi, H \psi \rangle \, &lt; \, 0. Then we can use a remarkable property of the r 2r^{-2} potential to show that HH is not bounded below. Namely, HH has a kind of symmetry under dilations. You can guess this by noting that both the Laplacian and r 2r^{-2} have units of 1/length 2{}^2. Indeed, if you take any smooth function ψ\psi, dilate it by a factor of α\alpha, and then apply HH, you get α 2\alpha^{-2} times what you get if you do these operations in the other order. This implies that if

ψ,Hψ=Eψ,ψ, \langle \psi, H \psi \rangle = E \langle \psi, \psi \rangle ,

we can dilate ψ\psi and get a function obeying the same equation with EE replaced by α 2E\alpha^{-2} E. Thus, as soon as EE can be negative, it can be made arbitrarily large and negative by choosing α\alpha to be very small. Thus HH is not bounded below.

Next, what is special about c=34c = \tfrac{3}{4}? This is more subtle. For any value of cc \in \mathbb{R} we can find spherically symmetric solutions of ( 2+cr 2)ψ=iψ ( -\nabla^2 + c r^{-2})\psi = i \psi on 3{0}\mathbb{R}^3 - \{0\} that are nonzero and smooth. When c<34c \, &lt; \, \tfrac{3}{4}, and only in this case, some of these solutions ψ\psi lie in L 2( 3)L^2(\mathbb{R}^3). This dooms the chance of HH being essentially self-adjoint, because it implies H*ψ=iψH\ast \psi = i \psi. If HH were essentially self-adjoint H*H\ast would be self-adjoint, and it is easy to see that a self-adjoint operator cannot have ii as an eigenvalue.

When c<34c \, &lt; \, \frac{3}{4} the operator HH has more than one self-adjoint extension from C 0 ( 3{0})C_0^\infty(\mathbb{R}^3 - \{0\}) to some larger domain. To classify these we can use separation of variables, writing 2\nabla^2 as a sum of a radial part and an angular part, assuming the angular dependence of ψ\psi is given by a spherical harmonic Y mY_{\ell m}, and doing a change of variables u=ψ/ru = \psi/r to reduce HH to the ordinary differential operator

d 2dr 2+(c+(+1))1r 2 - \frac{d^2}{d r^2} + \left(c + \ell(\ell+1)\right) \frac{1}{r^2}

on the half-line (0,)(0,\infty). We can completely classify self-adjoint extensions of this differential operators from C 0 (0,)C_0^\infty(0,\infty) to larger domains; the answer depends on cc and \ell. A choice of self-adjoint extension is a choice of boundary conditions at r=0r = 0, and this says how the phase of an incoming wave changes as it reflects off the origin and bounces back. Finally, we can assemble the results for different spherical harmonics to classify self-adjoint extensions of HH.

There exist many self-adjoint extensions of HH that respect the rotational symmetry of the inverse cube force law, but for c<14c \, &lt; \, -\tfrac{1}{4} the extension must break the dilation symmetry discussed above. This is what physicists call an ‘anomaly’: a symmetry of a classical system that fails to be a symmetry of the corresponding quantum system. Intriguingly, for some even lower values of cc one can choose a self-adjoint extension that is symmetrical under a discrete subgroup of dilations. Determining precisely which values these are seems to be an open problem.

To explore this topic thoroughly, I recommend first this:

then this:

and finally this:

The first is an excellent overview of problems associated to singular potentials, including the inverse cube force. The second delves into self-adjoint extensions of the ordinary differential operators mentioned above, and the third works them out with exquisite thoroughness.

April 16, 2026

Clifford JohnsonComputing Correlators

[A more technical post follows]

My most recent paper, out on the arXiv today, is very exciting to me because it seems to be a genuinely new way of computing some important quantities and it is devilishly simple. So simple that I worried for months that it is all super-obvious to everyone. But another voice within me said to myself: Well if it is so obvious, why has nobody published it? Another (paranoid) voice within said: Maybe someone has published this method, and I just can't find it in the literature...

Well, I decided that the best way to find out for sure is to put it on the arXiv and within a short time someone will email to say that I missed their important work. So, while I wait for that email (as I start writing it's only been 30 minutes since it has been "out there", so there's time), let me say a few things about why I like the many results in the paper.

I was already pleased enough with the core part of the paper that I was going to write a swift four-pager about it back in February. The core point being that I figured out how to build on work I'd done in a paper back in 2024 (expanded on with followup work I did with Wasif Ahmed and Krishan Saraswat, a student and postoc). Back in 2024, I found (here) a really nice way (almost miraculous in how it worked) of writing all the corrections to the spectral density of a class of models in terms of one function [latex]u_0[/latex] and its derivatives. It was obtainable from one simple ordinary differential equation (ODE) called the Gel'fand-Dikii equation, which takes in the function [latex]u_0(x)[/latex] as input. The ODE is for a special quantity called the diagonal resolvent [latex]{\widehat R}(x,E)[/latex]. You integrate that quantity [latex]\widehat R(x,E)[/latex] with respect to [latex]x[/latex] and you're more or less home. In general, it is a messy quantity that does not integrate to anything nice. But just when the function [latex]u(x)[/latex] obeys the "string equation" it is supposed to (as dictated by the governing model's physics), then [latex]{\widehat R}(x,E)[/latex] is a total derivative (a seeming miracle-see later), and the corrections it gives to the density become of just the right form!

Those corrections can be called [latex]W_{g,1}(E)[/latex] where the [latex]g[/latex] is the order in perturbation theory. [latex]g=0[/latex] is leading order, [latex]g=1[/latex] is the torus, [latex]g=2[/latex] the double torus, etc. Indeed [latex]g[/latex] is the number of handles or "genus" of an associated Riemann surface. The one subscript on the other hand, corresponds to the one energy entry available when just discussing the density [latex]\rho(E)[/latex]. All the [latex]W_{g,1}[/latex] end up being written nicely in terms of a function [latex]u_0(x)[/latex] and its derivatives, evaluated at a special point.

An already nice feature (among many) of the construction was that this one ODE, recursively solved, gave rise to the [latex]W_{g,1}[/latex] of many different problems across a range, including certain random matrix models, gravity problems, intersection theory and topology, and so on. All you need to do is change the function [latex]u_0(x)[/latex]. Moreover, for this (wide) class of problems, you can compute the desired results faster and with way less machninery than other methods, such as topological recursion, which was an interesting observation. This includes very famous problems like the Weil-Petersson volumes (of the compactified moduli space [latex]\overline{\cal M}_{g,1}[/latex] of Riemann surfaces with genus [latex]g[/latex] and [latex]n=1[/latex] boundaries) and generalisations. Another nice feature is that you also get non-perturbative data beyond the genus expansion, an aspect I explored recently (in this paper) with student Joao Rodrigues, and expert in resurgence techniques.

The core breakthrough of the new paper is this: For some time, I've wondered how to compute correlators for more energies (amounting to multi-point correlators of [latex]\rho[/latex]) in this same way: [...] Click to continue reading this post

The post Computing Correlators appeared first on Asymptotia.

April 13, 2026

John PreskillHow I learned to stop worrying and…no, I’ve always adored entropy

When I was pursuing a PhD at Caltech, so was my friend Jeremy. He used to throw a dinner party every few months. The email invitations welcomed friends to partake of his cooking and, if we wished, to help him cook. I didn’t help cook; but, when I arrived, the mess of pots and pans drew me to the kitchen like vinegar drawing a pathological fly. I couldn’t sit still while cookware needed cleaning, so I scrubbed and rinsed the pans and spoons and bowls. Jeremy, an applied-physics student, commented on my adeptness at decreasing entropy.

It’s the story of my life, I replied.

In fourth grade, my classmates and I cleaned our desks every Friday afternoon. Once a student finished, my teacher dismissed him or her onto the playground. My neighbor’s desk horrified me like the disaster in a hurricane’s wake, so I neatened his desk after finishing with mine.1 Another friend requested the same favor. A third classmate offered to pay me for cleaning his desk, but I’d have undertaken the chore for its own sake. Ordering the world offered me fulfillment.

From cleaning a fourth-grade desk, I progressed to pursuing a PhD in theoretical physics. The two pursuits might seem to resemble each other no more than Dr. Jekyll and Mr. Hyde; yet, to me, the path between them is but a step. I trained as a theoretical physicist because I love organizing ideas. Caltech paid me to build models, propose definitions and theorems, and structure proofs—to dream up ideas and identify the optimal arrangements for them. I needed that pay, being an adult, as I hadn’t needed my fourth-grade classmate’s desk-cleaning fee. Yet I organized ideas for the same reason that drove me to organize my neighbor’s notebooks.

Many people have called entropy a measure of disorder. To see why, imagine that Jeremy’s crew has used thirty utensils while cooking. The chefs can have scattered the utensils across the kitchen in many ways: they may have dropped forks on the floor, left spoons in the sink, arranged spatulas on the drying rack, or filled a vase with knives like a modern-art bouquet. In few of these configurations do the forks lie in their compartment of the utensil drawer, the spoons lie in their compartment, etc. We call such configurations neat. Most of the other configurations, we call messy. 

A system’s entropy is the number of configurations consistent with known large-scale properties of the system, such as the number of forks.2 More configurations are consistent with messiness (and a fixed number of forks and so on) than with neatness (and the same number of forks and so on). Messiness tends to correlate with high entropy. People often say, therefore, that entropy quantifies messiness. Hence Jeremy’s complimenting me on my decreasing of entropy.

Jeremy’s dinner parties came to mind as I read the book The Mattering Instinct, published by Rebecca Newberger Goldstein this January. Rebecca is a philosopher of science and a writer. I had the good fortune to meet her through my undergraduate mentor Marcelo Gleiser, who’s had another cameo or two on Quantum Frontiers. Rebecca’s latest book covers what she calls the mattering instinct: the longing to know that we matter. 

We spend scads of energy and time on securing our “survival and flourishing,” as Rebecca says. We feed ourselves; work to earn money to purchase food; clean, shelter, and clothe ourselves; ingrain ourselves in societies that offer some degree of security; and more. Do we deserve all this effort? We long for assurance that, in the immortal words of L’Oreal, we’re worth it. 

Survival and flourishing, Rebecca writes, requires us to decrease entropy. Every closed, isolated system’s entropy increases or remains constant, according to the second law of thermodynamics. Entropy increases as a system becomes more uniform, loosely speaking. The system’s particles spread out across space, these particles’ temperature comes to equal those particles’ temperature, and so on. In contrast, your body exists because its particles clump together in a certain shape consistently. You withstand heat waves and snow because homeostasis maintains your temperature despite your environment’s temperature. You keep your body’s entropy low to survive. Rebecca therefore casts us as fighting entropy.

As a thermodynamicist, I agree with Rebecca. Yet I also adore entropy. It helps explain why time flows, quantifies uncertainty, and determines the maximal efficiencies with which we can perform tasks such as communication. What versatility and richness! Entropy also embodies tension and subtlety: its mathematical definition looks obscure at first glance, yet entropy helps explain familiar phenomena such as aging. For these reasons, before beginning my PhD, I told a potential advisor that I could imagine devoting the next five years of my life to entropy.

I therefore aspire to rehabilitate entropy’s reputation. Novelist Terry Pratchett endeared mortality to millions of readers through anthropomorphism. His character Death, a mainstay of the Discworld series of novels, elicits empathy and fondness. I won’t anthropomorphize entropy here,3 but I aim to replace conflict with cooperation in the narrative above. To survive and flourish, I hold, we partner with entropy. How? We create oodles of entropy in our environments. This entropy increase offsets the entropy decrease that supports life.

For example, imagine working at a desalination plant. You’d process high-entropy water throughout which salt has spread. You’d concentrate the salt in a tiny region, reducing the water’s entropy. This reduction, producing fresh water, could support your city’s drinking, cooking, and toothbrushing needs.

To reduce the water’s entropy, you’d create loads more entropy. You’d eat breakfast before work, consuming energy stored neatly in your waffle’s chemical bonds. Your body would later break the bonds, releasing the energy. Some energy would power your muscles, so you could program the desalination system, test its output, etc. But much of the chemical energy would transform into heat radiated by your body. The heat would warm up the air molecules around you, magnifying their random jigglings and jostlings. You’d increase the entropy of the air—your environment—to decrease the water’s entropy. The air’s entropy increase would outweigh the water’s entropy decrease.

Organisms survive and flourish by producing entropy in their environments. In fact, organisms have a knack for generating entropy. Entropy and life thereby further each other. A glass-half-full thinker could conclude that we partner with entropy.

So did I partner with entropy as a PhD student, applying it to solve problems in quantum information theory and thermodynamics. So did I partner with entropy in fourth grade and at Jeremy’s apartment, deriving satisfaction from my cleaning. Rebecca would call these activities’ ultimate aim (beyond the aim of, e.g., not sitting beside a pigsty in fourth grade) mattering. She writes that we reduce entropy (within our immediate vicinities) to satisfy the mattering instinct. Rebecca’s proposition describes my behaviors with uncanny precision, I realized upon reading her book.

Which I’ve now finished. So pardon me while I return to washing forks in the quantum kitchen of the universe. 

With thanks to Jeremy for his friendship…and food.

1I also ensured that my neighbor brought home, every afternoon, the sweater he’d brought to school that morning. Before I took charge, he’d ended up with three forgotten sweaters crammed into his cubby.

2At least, one entropy is. Many other entropies exist.

3If you anthropomorphize entropy elsewhere, let me know.

March 26, 2026

Tim GowersGroup and semigroup puzzles and a possible Polymath project

An Artin-Tits group is a group with a finite set of generators a_1,\dots a_k in which every relation is of the form (ab)^r=(ba)^r or (ab)^ra=(ba)^rb for some positive integer r, where a and b are two of the generators. In particular, the commutation relation ab=ba is allowed (it is the case r=1 of the first type of relation) and so is the braid relation aba=bab (it is the case r=1 of the second type of relation). This means that Artin-Tits groups include free groups, free Abelian groups, and braid groups: for example, the braid group on k+1 strands has a presentation with generators a_1,\dots,a_k, where a_i represents a twist of the ith and (i+1)st strands, and relations a_ia_j=a_ja_i if |i-j|>1 and a_ia_{i+1}a_i=a_{i+1}a_ia_{i+1}.

A few weeks ago, I asked ChatGPT for a simple example of a word problem for groups or semigroups that was not known to be decidable and also not known to be undecidable. It turns out that the word problem for many Artin-Tits groups comes into that category: the simplest example where the status is not known is the group with four generators a,b,c and d where c and d commute and all other pairs of generators satisfy the braid relation xyx=yxy.

My interest in this was initially that I was looking for a toy model from which one could learn something about how mathematicians judge theorems to be interesting. I started with a remarkable semigroup discovered by G. S. Tseytin that has five generators, seven very simple relations, and an undecidable word problem. From that I created a puzzle game that you might be interested to play. (NB the puzzles were not showing up in the public version, but that should be sorted out now.) Each puzzle is equivalent to an instance of the word problem for Tseytin’s semigroup, but the interface makes it much more convenient to change words using the relations than it would be to do it with pen and paper.

I was hoping that because every mathematical problem can in principle be encoded as a puzzle in this game, one might be able to build a sort of “alien mathematics”, where a theorem was an equality between words, and a definition was a decision to introduce a new (and redundant) generator g, together with a relation of the form g = w, where w is a word in the current alphabet. Theorems would be particularly interesting if they were equalities between words that could be established only by a chain of equalities that went through much longer words, and definitions would be useful if the new generators satisfied particularly concise relations (which would allow one to build “theories” within the system). I still hope to find a word problem that will allow such a project to take off, but in the end, after a lot of playing around with the game linked to above, I have decided that Tseytin’s semigroup is not suitable. The reason is that it is designed so that arbitrary instances of the word problem for groups can be encoded as word problems in this semigroup, and once one gets used to the game, one starts to see how that encoding can work. Furthermore, one seems to be driven towards the encoding — I don’t get the impression that there’s a whole other region of this semigroup to explore that has nothing to do with the kinds of words that come up in the encoding. And if that impression is correct, then one might as well start with the word problem for some group in unencoded form, or alternatively look for another semigroup. Nevertheless, I find the Tseytin game quite enjoyable: I won’t say more about it here but have written a fairly comprehensive tutorial that you can open up and read if you follow the link above.

This is perhaps the moment to say that the words “I created a puzzle game” are slightly misleading. For one thing, I discussed the idea of gamifying Tseytin’s semigroup about three years ago with Mirek Olšák, a former member of my automatic theorem proving group, and he created a basic prototype in Python. But the main point is that I do not have the programming skills to create a game that can be played in a web browser — I vibecoded it using ChatGPT.

After that experience, I thought that maybe I would have better luck if instead of looking for a group or semigroup with undecidable word problem, which might well have been explicitly designed with some encoding in mind, I looked for a word problem for which the decidability status was unknown. That way, it wouldn’t have been designed to be undecidable, but might nevertheless just happen to be undecidable and provide a nice playground of the kind I was (and still am) after. And that is what led me to the Artin-Tits groups.

However, those don’t seem to be suitable either, because it is conjectured that they all have decidable word problems. I have created a game for the Artin-Tits group mentioned above, which you can also play if you want, but I have found it very hard to create interesting puzzles. (NB again there was a problem with the puzzles showing up, with only a tiny handful being there, but that should now be sorted out.) That is, I found it difficult to find words that are equivalent to the identity but that are not easily shown to be equivalent to the identity. One nice example comes from the fact that the subgroup generated by a,c and d is isomorphic to the braid group B_4 with four strands. The late Patrick Dehornoy found a very nice example of a braid with four strands that is equal to the identity but not in a completely trivial way. A picture of it can be found on page 4 of this paper.

This is where a potential Polymath project comes in. An initial goal would be to determine the decidability status of this one small Artin-Tits group: if we managed that, then we could consider the more general problem. And the way I envisage approaching this initial goal is an iterative process that runs as follows.

  1. Devise an algorithm A_0 for solving word problems in the group.
  2. If the current algorithm is A_i, then search for a puzzle P that A_i fails to solve.
  3. If the search is successful, then devise an algorithm A_{i+1} that solves all the puzzles that A_i solves, and also solves P. Make A_{i+1} the current algorithm and return to step 2. Devising A_{i+1} can be done by playing the game with puzzle P several times until one has the feeling that one has solved it in a systematic way.
  4. If the search is unsuccessful and the current algorithm is A, then attempt to prove inductively that A always succeeds: that is, to prove that if A solves P and Q is obtained from P by applying a single relation in the group, then A solves Q.
  5. (Maybe) If the inductive step doesn’t seem to work, then try to use that failure to devise a puzzle that defeats A and return to step 3.

I have done a couple of iterations already, with the result that I now have an algorithm — let’s call it A_0 — that solves (basically instantaneously) every puzzle I throw at it, including the one derived from Dehornoy’s braid mentioned above. So a subgoal of the initial goal is to find a puzzle that has a solution but that A_0 doesn’t solve. If we can’t, then maybe it is worth trying to prove that A_0 solves the word problem for this group. One comment about the algorithm A_0 is that it never increases the length of a word, though it often does moves that preserve the length. I was told by ChatGPT that it is an unsolved problem whether every braid that is equal to the identity can be reduced to the identity without ever increasing the number of crossings. Of course that isn’t a guarantee that it really is unsolved, but if it is, then that’s another interesting problem. It also means that I would be extremely interested if someone found an example of a word in the Artin-Tits group that can be reduced to the identity but only if one starts by lengthening it.

I’m hoping that this will be an enjoyable project for people who like both mathematics and programming. On the maths side, there is an unsolved problem to think about, and on the programming side there are lots of possibilities: for example, one could write programs that explore the space of words equal to the identity, trying to do so in such a way that there is a reasonable chance of reaching a word where it isn’t obvious how to reverse all the steps and get back to the identity. The game comes with a “sandbox” that has a few tools for generating words, but at the moment it is fairly primitive, and I would welcome suggestions for how to improve it.

It seems to me that as Polymath projects go, one can’t really lose: either we find an algorithm for the word problem, which establishes an unknown (at least according to ChatGPT) case of a decidability problem, or we find a suite of harder and harder puzzles, creating a more and more challenging and entertaining game and obtaining a deeper understanding of the Artin-Tits group in the process.

This Artin-Tits game has only a rather rudimentary and not very good tutorial created by ChatGPT. It’s easy to play once you’ve got the hang of it, and in the end I think the easiest way to get the hang of it is to watch someone else play it for a couple of minutes. I have therefore created a video tutorial, which you can find here. The video itself lasts about 25 minutes, but the tutorial part is under 10 minutes: what makes the video longer is an explanation of various features for making it easier to create puzzles, and also an explanation of how the algorithm works, which again is much easier to explain if I can demonstrate it on screen than if I have to write down some text. (I do have some text, since that is what I used as a prompt for ChatGPT to implement the algorithm, but it took a few iterations to get it working properly, so now I’m not sure what the quality of the code will be like, or even whether it is doing exactly what I want, though it appears to be.) Although I have embedded the video into this post, if you actually want to see what is going on you will probably need to watch it on YouTube using the full screen. I recommend not watching the explanation of the algorithm until you have played the game a few times first. Also, I should warn you that the games in the Advanced category are not all soluble, but games 1, 5 and 6 definitely are. Another thing I forgot to say on the video is that if you want you can “rotate” a word by clicking on one of its letters and dragging it to the right or left. If the word is equal to the identity, then its rotations will be as well, so this is a valid move. There is also a button for disabling this rotation facility if you want the puzzle to be slightly more challenging.

A quick note about the availability of the two games. They are hosted on the Netlify platform. I am on their free plan, which gives me a certain number of “credits” each month. I’m not sure how quickly these will run out, largely because I have no idea how many people will actually think it interesting to play the games. If they do run out, then the games will cease to be available until the credits are renewed, which for me happens on the 10th of each month. If this has happened and you are keen to play one of the games, another option is to download the html files and open them in your browser. Here is a link to the Artin-Tits game and here is one to the Tseytin game. If you are feeling particularly public-spirited, especially if you think you will play quite a lot, then you might consider doing that anyway, so that the Netlify credits run down more slowly. If the running out of the credits is quick enough for all this to be a real issue, then I may move to a paid plan.

I’ll finish with a quick tip for playing the Artin-Tits game, which I mentioned on the video but perhaps didn’t stress enough. Many of the moves consist in selecting three consecutive letters and using a relation of the form xyx=yxy to change them. Easy examples of this are replacement rules such as aba\to bab. But what if inverses are involved? I’ll represent inverses of generators with upper-case letters, so for example abA represents aba^{-1}, which in the game would be a white a followed by a white b followed by a black a. The word abA turns out to equal Bab. To remember this, a simple rule is that two letters of the same colour can be bracketed together and “pushed past” the third letter, which retains its colour but changes its value. Here, for example, we write abA as (ab)A and then swap them over, changing A into B in the process, getting B(ab), or Bab without the brackets. In group theoretic terms, this is of course saying that A and B are conjugates, and that ab is what conjugates one to the other. But when playing the game it is convenient to remember it by thinking that when you see a subword such as bDB, you can push the DB (and in particular the D) to the left, getting DBd.

February 15, 2026

David Hoggnumber of exposures per visit?

I wrote code this weekend to look at the question of how we should visit a star in the upcoming Terra Hunting Experiment. The current (straw-person) plan is that we will observe each visible star once per night for ten years, with one exposure of a sensibly-chosen exposure time at each visit. Is this a good idea? I was interested in this problem for two reasons. The first is that binning is sinning, with the corollary that bigger bins are worse than finer bins, and a single, long exposure is a very big bin. The second reason is that when there are non-trivial noise sources (like the quasi-periodic variations from p-mode oscillations of the surfaces of Sun-like stars), a few negatively- (or interestingly-) correlated noise draws can be combined in ways that are substantially more informative than by taking the average.

Of course, if you split an exposure (with a standard CCD, say) into sub-exposures, you take on real costs: There is a read time, which is time you aren't integrating, and there is a read noise, that affects each new exposure. So the best strategies are a complicated function of read time, read noise, and the signal-to-noise at which the stellar p-mode oscillations are visible in any realistic data. Related: There are amazingly different and interesting strategies with up-the-ramp detectors that are used in the infrared.

One final comment is that the objective, in my strongly held view, is to optimize the amount of information (about, say, the center-of-mass radial-velocity changes of the target star) per unit wall-clock time. We are paying for wall-clock time; let's get as much as we can out of it.

February 13, 2026

David HoggThe LLMs, and why do we do astrophysics?

Today my rant on LLMs and the practices of our field hit the arXiv. I was scared to post it, because it is such a weird contribution, and it is so revealing about myself and my own political positions and hangups. But I have to say: I got great and supportive feedback all day.

I got two comments on saying ACAB in the literature. The Astronomer Royal of Scotland quoted (on BlueSky) the last sentence, which I put there because Andy Casey (Monash, Flatiron) insisted. Many people sent me appreciation and thank-yous, and many people sent me comments and objections. Always constructive. The whole experience made me feel very happy about the state of our field and the way we all interact. I think maybe there will be critical mass to write some kind of collection of essays on the subject. That's a plan for 2026.

February 05, 2026

Matt Strassler The Physicists and Mr. Epstein

Mr. Epstein was not only a world-class child abuser, he was a big fan of theoretical high-energy physics and of theoretical physicists. Some of my colleagues, unfortunately, got to know him. A number who were famous and/or had John Brockman as a book agent were even invited to a physics conference on Epstein’s private island, well before he was first arrested. This was no secret; as I recall, a lot of us heard about the existence of this conference/trip, but we hadn’t heard Epstein’s name before and didn’t pay much attention (ho hum, just another weird billionaire).

Personally, I feel quite lucky. The Brockman agency rejected the proposal for my recent book without comment (thank you!); and my research is mostly considered unimportant by the Brian Greenes of the world. As a result, I was not invited to Epstein’s island, never made his acquaintance, and blissfully avoided the entire affair. Clearly there are some benefits to being considered ordinary. And so — I’m sorry/not-sorry to say — I can’t tell you much about Epstein at all, or about how certain physicists did and did not interact with him. Regarding my colleagues who did get to know him, I can’t speak for them, since I wasn’t there, and I don’t know to what extent Epstein hid his immoral activities when they were around. It’s up to them to tell their own stories if they feel the need to do so (and I hope a couple of them do, just to clear the air.) Personally I tend to give them the benefit of the doubt — probably some literally didn’t know what was up until Epstein’s arrest in 2008, while perhaps others felt there wasn’t much they could do about Epstein’s actions on his own private island. I imagine they are deeply embarrassed to have been caught in this horrible man’s ugly web.

Fans of physics come in all shapes and sizes, and some have large wallets, large egos, and/or large ambitions. Among the wealthy supporters, we can count Alfred Nobel himself; billionaires sit on important scientific institute and university boards, and the more recent Breakthrough Prizes were funded by deep pockets. The extreme wealthy have outsized influence in our country and in our world, and one could argue that their influence in 2025 was not for the better. Usually, though, the influence in physics and related fields tends to be relatively benign, funding postdoctoral researchers and graduate students who deeply want to do science but also need to eat. That said, sometimes donors fund non-essential fields at the expense of critical ones, or favor theoretical research over the gathering of crucial experimental data, or push money on famous rich organizations when there are poor ones that are equally deserving and far more needy.

When gazillionaires, on their own initiative, come calling on non-profit organizations, whether they be community centers, arts organizations, or universities, they pose a problem. On the one hand, it is the job of anyone in a non-profit organization to help raise money — fail to do that, and your organization will close. When a single person offers to permanently change the future of your program, you would be derelict in your duty if you did not consider that offer. On the other hand, donors who might have ethical or criminal problems could drag the organization’s name through the mud. Worse, they might be able to force the organization itself to do something ethically questionable or even illegal.

There is a clear lesson for young academics and other up-and-coming non-profit actors in the Epstein affair: the more money potentially offered to our organizations, the more carefully we must tread. Money is power; power corrupts; and every pursuit of dollars, even for the best causes, risks infection. We can’t be large-scale non-profit fundraisers without doing serious and thorough background checks of the biggest donors; we have to question motives, and we can’t look the other way when something seems amiss. Those of us with clear hearts and honest pursuits tend to assume the best in other people. But we have to beware of those hoping to bolster their reputations, or clean their consciences, by giving away “generously” what they never deserved to have.

December 01, 2025

Secret Blogging SeminarCongress proposes cutting of all funding to US academics who mentor Chinese students

I’m writing to point out a potential law which should be gathering more opposition and attention in math academia: The Securing American Funding and Expertise from Adversarial Research Exploitation Act. This is an amendment to the 2026 National Defense Authorization Act which has passed the House and could be added to the final version of the bill during reconcilliation in the Senate. I’m pulling most of my information from an article in Science.

This act would ban any US scientist from receiving federal funding if they have, within the last five years, worked with anyone from China, Russia, Iran or North Korea, where “worked with” includes joint research, co-authorship on papers, or advising a foreign graduate student or postdoctoral fellow. As I said in my message to my senators, this is everyone. Every mathematician has advised Chinese graduate students or collaborated with Chinese mathematicians, because China is integrated into the academic world and is one fifth of the earth.

This obviously isn’t secret, since you can read about it in Science, but I am surprised that I haven’t heard more alarm. Obvious people to contact are your senators and your representatives. I would also suggest contacting members of the Senate armed services committee, who are in charge of reconciling the House and Senate versions of the bill.