Planet Musings

October 01, 2026

Proofs and Prompts — A new type of paper mill

Eric Dolores-Cuenca, Postdoc at Yonsei University

The purpose of this letter is to warn the mathematical community of a new type of paper mill that could potentially appear as a consequence of the advances in artificial intelligence.  I am a postdoc working on Math and AI, and I organize an international online seminar on math and machine learning.

Firstly, there exist groups that produce fake research papers and sell them to scientists; these groups are commonly called paper mills. According to this study, at least 2% of all research submissions in all sciences are submitted by paper mills.

Secondly, there also exist start-ups, companies and universities working on solving mathematical problems with machine learning. In particular, some use tools that take information from a mathematical problem and return, when possible, experiments to prove it. The level of human interaction varies: in some cases the user has to provide only a text description of the problem; while for other programs, the user must write code that evaluates the fitness of the solutions proposed by agents.1  AI agents are also capable of transforming the output of those experiments into a written research paper. While it is not clear how much interaction was needed to produce the preprint about the Navier-Stokes problem or to solve the Jacobian conjecture, in this post I want to give emphasis on the possibility of an artificial intelligence solving completely a problem only using a prompt as input.

With these two antecedents, I would like to ask the reader: what would happen if a group trained an algorithm to recognize important mathematical questions? The algorithm can work in parallel and mass produce meaningful questions. If we now give this list of questions as inputs to algorithms that try to solve them, and we furthermore assume that they manage to solve a small percentage of the input questions, then the final products are research papers that are publishable since they have solved meaningful questions for mathematicians. Finally, what if the company sells those finished results to mathematicians? This would create a paper mill that produces correct papers and sells the authorship. Moreover, if an AI that solves problems has a certain level of quality, an AI that formalizes the proof can further filter the papers.

One can wonder what the danger is if paper mills are producing knowledge, that is, publishable papers. Fraudulent credit is one of the problems; affecting hiring, promotion, grants, and reducing the credibility of the publication process.

In informal conversations, I asked some of the experts who attended the Mathematics × AI: Challenges and Opportunities event at the Korea Institute for Advanced Study and the International Conference on Machine Learning 2026 about the possibility of this new type of paper mill. Some of the people I talked to are collaborating with AI companies, while others were participants with different levels of experience in AI.

While I personally expected to be assured of the impossibility of this new type of paper mill, some of the experts considered this scenario plausible within a few years, and they could not think of a solution for journals to detect this problem. Some told me they have students working on an algorithm to predict important problems. I was encouraged by them to keep talking about this problem in order to raise awareness, which led me to write this letter.

Why would this be a very difficult problem for journals?

From the point of view of the journals, there are no steps in the review process to ensure that the person who submitted the paper did not buy the paper from a paper mill, or to certify that the person is aware of the content of the paper. Note also that using a detector to see if a paper is written by AI is useless: a person can buy finished papers and rewrite them.

We are now seeing talks in which people used AI in their work. Note that the quality of the talk may not be related to the use of AI. The speaker could be afraid of public speaking, perhaps the speaker does not have enough experience giving talks, faces a language barrier, or is new to the topic and knows just the minimum required for the result to hold.

I call on the community to discuss what changes journals would need to implement if we ever reach that level of technology.

Acknowledgement: I would like to thank Susana Lopez Moreno for related discussions.

  1. The following examples, listed alphabetically, are not exhaustive: Aletheia, AlphaEvolve, Astra, Iteris, OpenEvolve, Rethlas, and ShinkaEvolve. ↩︎

Received 11 September 2026.

Proofs and Prompts — What about the preprints?

Manuel Rissel, tenure-track faculty at ShanghaiTech University

With the onset of “LLM driven math turbulence”, various guiding principles that seemed for granted are being shaken up a little. I can imagine that today someone could upload a new preprint introducing a rich theory that prepares all kinds of exciting unknown terrain for future exploration, and on the next day, before a word was talked about that theory with anyone, LLM driven agents may already have analyzed and used material from that preprint in the process of generating proofs to a list of conjectures which someone else (or some company) requested from a chatbot. Whether this would be perceived as good or bad likely depends on who one asks under which circumstances, but it seems natural to wonder: would the only guaranteed way for ruling out such a scenario (if one wanted to) consist of not uploading new research as preprints? And,

would a majority of mathematicians like their preprints (including the LaTeX files) being crawled by arbitrary automated robots?

There is hardly one answer fitting all or a majority. For instance, researchers may enjoy seeing their not yet peer-reviewed proofs appear in LLM generated answers. Some may want to actively support LLM driven research or at least have it stabilized at an equilibrium as quickly as possible. Others may not want to contribute new ideas immediately to automated LLM driven research or to external companies with other interests. Also, there can be nuances, as some may be fine with sharing titles and abstracts of preprints with robots, but would like to reserve the proofs mostly for mathematicians during the peer-review stage.

Nevertheless, there appear not many options on the table, as the preprint server mostly used by the math community operates based on principles such as the open distribution of knowledge, and openness is something that hardly anyone would like to change. But what means open? How to choose the topology that’s best for sustained mathematical science development? Is there a natural mechanism that ensures sufficient openness of preliminary research findings but still provides a new lever for the community to influence slightly the balance between the amount of digested mathematics and the number of math manuscripts?

One hypothetical thought could be that some authors may choose to encrypt their freshly indexed preprints, keeping the title and a possibly long abstract available to the public, with automated decryption of the main body to occur after a well-defined embargo period, e.g. 1-2 years (or any time scale larger than that of LLM development). Such a preprint would be compiled by the server after upload and then encrypted, retaining a readable copy only for private access or moderation. It would be time stamped and versioned, just like done by the arXiv right now, hence the authors could without worries start talking to others about the novelties of their new theory and send readable versions to journals and colleagues. To prevent meaningless encrypted preprints, there could be requirements for using this particular service: e.g., a limit of encrypted uploads per verified researcher per year. Other researchers interested in learning about the details contained in an encrypted preprint could invite the authors to visit them and give a talk, invite them to online seminars, engage in private exchange…\ldots or to take this thought ad absurdum, someone could prompt an LLM…\ldots

The drawbacks of voluntary encryption of preprints may outweigh the positives, but it could still be valuable to debate about the role of preprints in a changing research environment. Our way of handling preprints potentially influences peer-review, knowledge digestion, and the ability of external forces to act on our profession.


Received 8 September 2026.

September 30, 2026

Proofs and Prompts — Open letter to the European Mathematical Society

Snorre H. Christiansen, Professor at the University of Oslo

The EMS should establish a structure to shepherd the use of artificial intelligence in mathematics research. First version sent to the EMS on 20 September 2026.

A European institute for AI enhanced research in mathematics

This open letter is a response to the latest developments in AI (including contributions to the Millennium problem on Navier-Stokes) as they impact the mathematical community.

My impression is that many research mathematicians would feel a loss of meaning if the most advanced research were to consist mainly in efficient prompting and agent management, especially if these tools are outside of our control. This sense of loss compounds with scepticism about where the broader society is driven, by the interests of a handful privately owned outsized AI companies. What’s more, in most cases these companies are subject to politics and jurisdictions that we can hardly count on swaying.

I believe that the mathematical community is well placed to develop its own AI models, to regain control over its destiny. The community has competent and motivated manpower, especially if research as we knew it is upended, as seems to be already the case. I am sure that mathematicians can develop the AIs that are best at doing maths and that they are best placed to help regulate the use of this tool, notably in regards to the education and recruitment of future generations of mathematicians. Mathematical thinking is as needed now as it has ever been.

This communal AI could take the form of an institute owned by EU/EEA through the European Research Council (ERC), governed by the EMS and partnered with promising European AI companies for mutual benefit (I have noted that Mistral has developed Leanstral).

It would be subject to European law, contribute to European sovereignty and would give European mathematicians a hand on the steering wheel in matters that are important to them. Communal values have been expressed in particular in the Leiden declaration. The declaration by Fields medalists relates our concerns to the ones expressed in other arts and sciences. The initiative should be mindful of the solidarities felt by mathematicians across continents.

Mathematicians have always had a broad spectrum of views on matters of ethics, especially regarding military technology. I suppose this will also be the case regarding AI. I understand that some will prefer to stay away from AI all together, and that’s fine. Protecting human agency in mathematics can take many forms.

This initiative would provide a structure that gives willing mathematicians a possibility to defend their values and regain a sense of control over their destiny. If prompting is the future, I for one would like to prompt inside a structure I can support. This way we can at least try to align future AIs with the interests and sense of dignity we defend as mathematicians and as human beings united by the Universal Declaration of Human Rights.

Traditional ways of doing maths would coexist with AI enhanced mathematical research developed in house and whose findings are shared on a need to know basis and published according to consensual protocols as collective endeavors (high energy physicists working at CERN have more experience with this).

I would also suggest that the EMS establish a Millennium type prize related to AI. It could be for instance to explain mathematically why deep neural networks work so well on some problems otherwise considered intractable (due for instance to the curse of dimensionality). This would give the EMS an occasion to communicate to the greater public about the role of mathematics in modern technologies.

Yours sincerely,

Snorre H. Christiansen


Received on 22 September 2026.

Terence Tao — Two reports

A brief post to note two reports that just came out:

Proofs and Prompts — The End of a Beautiful Era

Bogdan Chornomaz, postdoc at the Technion

As more and more central problems are nuked into oblivion by generated proofs, against the spreading AI-gloom the voices rise, emphasizing the internal value of deep mathematical thought and advocating for measures ranging from exercising extreme discretion with generative tools to seizing railway stations and the telegraph (or was it departments and classes of mathematical problems?) from the tightening grasp of the corporations.

This is understandable. Even if these pieces of advice seem hardly practical, we all need some venting as our dear field erodes beneath our feet, leaving soulless proof certificates where once the pure magic flowed. For what compels us to this calling, rather suboptimal in terms of earthly goals, if not its personal appeal, the sacred light of understanding, rising from hastily written notes in a dimly-lit room.

This intimate link is worth fighting for, perhaps by locking yourself in a cabin in the woods for the purposes of practicing pure thought (don’t try this at home; not without secured tenure, or at least a decent financial cushion). This way, no matter how much math were to succumb to the implacable march of AI, its spirit will be preserved.

It won’t.

No matter how elevated the guiding idea of a social enterprise, its body consists of boring stuff, such as career trajectories and job security, not necessarily reduced to, but surely resting upon the utility of the said enterprise. With the utility gone, the bottom to middle layers will be chipped away, their inhabitants turned into statisticians or worse.1 A masterful artisan, descending to the market for mass-produced furniture with the true understanding of a bedside table, will find little appreciation of his work.

Oh, come to me, math jobs market analyst, with hard data in hand; come and prove me wrong. But the September night is already offended by the sticky feeling of an exodus that has begun. An exodus that probably will be left unnoticed outside of the bubble, for the “I didn’t see you guys all the way over there” meme cuts both ways.

It is indeed sad. The utility I’ve been referring to is not necessarily, so to speak, utilitarian. Arguably, the authority of math is largely symbolic, grounded in its ability to find beauty emerging from triviality, poetry disguised as science. Those are smart guys, doing deep things, with integrity and a sense of purpose. I want to be a part of it. I don’t want to see it gone.

But even if AI progress in math were to come to an immediate halt, the high claim to possessing a unique and revealing insight into the nature of things would be shaken. The halo dims, and with it will dwindle the public attention and the funds. Maybe these smart guys will peacefully transition into areas with a more measurable impact; maybe they won’t. Maybe there won’t be anything to transition to.

Amidst a relatively mild disaster, as history turns its back on this land, I tell my young son: Learn mathematics, learn your French, learn your history.

Acknowledgements. Part of the blame for this text is on Yuval Filmus, who has drawn the author’s unhealthy attention to this resource.


Received 6 September 2026.

  1. This reference to Terence Tao’s quote from his interview, “On Math Education and Research”, is meant as a joke; no statistician was harmed in making this text. The full quote reads: “Most students who take math classes aren’t going to be mathematicians. They’re going to be engineers, statisticians – in many ways, that’s the more important mission of math education.” ↩︎

Terence Tao — How AI does, and does not, change the way I do math

[This is a guest post by Rachel Webb. This blog post was initially written in a different file format and converted using AI. — T.]

The way I do math experienced a great upheaval once before, halfway through my PhD. I was in a seminar with a group of graduate students and faculty developing new discipline-specific writing courses. The facilitator asked the question,

What does it mean to write well in your discipline?

Then, as now, I use writing as a proxy for doing. As a mid-career grad student, I had a quick answer to this question: good mathematical writing, hence good mathematics, is clear and correct. I was surprised at the dissimilarity between my answer and the answer of the other seminarians (who were more experienced but also other-disciplined): good writing says something interesting.

This discussion unlocked for me the meaning of doing research. Doing research is searching for something interesting to say. (Of course, once you find it, you have to say it in a way that is clear and correct.) In mathematics, the search for something interesting is often driven by our search for understanding, as noted by Thurston and endorsed by many since. But I think the math research process has a creative aspect that is not completely captured by this search. By contrast, the task of making something interesting to say is like writing a novel whose characters are at once subject to constraints of human experience but also presented in a way that comments on human experience. The characters in mathematical novels are definitions, the story text is the lemmas and theorems; these are constrained by truth, but can also be chosen strategically or artistically to capture certain facets of truth.

This is, abstractly, how I currently approach math research: I seek to understand what is happening at a fundamental level, yes, but I don’t think there is a unique way to understand. I also seek to take the things I do understand—inevitably, these are just a subset of the phenomena I would like to understand—and craft a narrative from them that is beautiful and interesting to my fellow mathematicians. Concretely, I find that “interest” in a mathematical context often derives from applications, either to the real world, or to other math. It also derives from proximity to high-profile open problems that serve as centers of mathematical conversations.

I don’t see LLMs as changing that approach much, but I expect they will drastically affect how I execute it.

The execution changes because now I have access to a machine that has read all the books and knows how many standard lemmas go. This speeds up the research process immensely and turns some of my lands of mathematical fantasy into worlds I can realistically start exploring (dream bigger dreams, says Antieau). Of course, taking the interstate instead of the side roads has its tradeoffs, but for any given leg of the journey I can choose which route to take. I will return to this idea in a moment.

My approach to research does not change because critically, I’m not convinced that the advent of LLMs changes my metric for mathematical “interest.” If we have historically regarded a paper as interesting if it comments in some way on a Millennium (or similar) problem, need we lessen our interest if the problem has a solution? Certainly there are many solutions, even many shards of understanding that do not constitute full solutions. These are all very interesting. We might know how to get to the red city, but we can continue to map out the surrounding terrain, now all the more valuable for the economic and strategic opportunities arising from metropolitan proximity.

I have said something about how I will continue to do math research, but I have not discussed why: what reasons do humans have to do math, if LLMs are capable (hypothetically, say) of producing math that is even more interesting than the math we can create? I believe that there are economic reasons, but I will not discuss those (both Sahai and Strogatz-Townsend have some thoughts). Instead I present two humanistic reasons. Neither of these is unique to math, just as math is not the unique human practice whose execution is affected by the advent of AI.

The first reason for humans to do math is that math is interesting to us individually. I enjoy doing math, so I will keep doing it, even if machines are better at it. There is a threat that LLMs will take the fun out of doing math by tempting us towards knowing the answer over understanding the solution. The temptation must be resisted. I should use AI only in ways that make math research more enjoyable for me and allow me to get understanding along with my wisdom (Proverbs 4:7). This may require experimentation. In fact, I and others already make analogous choices to use technology only when it is helpful to us personally and not every time it is economically “correct.” For example, I persuade my kids every week to walk seven miles round trip to church, even though we could drive. The commute costs well over two hours, but the increase in understanding between family members from the shared time and suffering is worth it.

The second reason for humans to do math is that math holds shared interest for multiple people at once, and in this way it creates communities. In some sense, this is what the other-disciplined seminarians meant when they said good writing is interesting: they meant that good writing is interesting to other people and hence is a piece of a larger conversation. Math research is a tool for drawing people together, in student-teacher relationships, in collaborations, at conferences, and at department colloquia and tea-times. Again, AI could weaken these social bonds by making (fear of) scooping more common, or just by making it easier to ask a machine than a colleague. But could does not automatically imply will. To quote Wendell Berry, may the age of AI be the age of knowing our mathematical neighbor. “Friends, every day do something that won’t compute.”

September 29, 2026

Scott Aaronson My “Knowmads” podcast on science and AI

Or click here if the above doesn’t work.

Recorded in-person in my office at UT Austin, with a bulleted list containing “ARC,” “Scalable Oversight,” and “Models” behind me on my blackboard for some reason (I no longer remember who put those there or why). 90 minutes long. Sometimes you see my disembodied arm waving in midair because of the way the cameras are combined. As always, I strongly recommend 2x speed for the correct experience.

This might actually be one of my best podcasts ever, although I wasn’t planning on that! Thanks so much to Bhavay Tyagi and Prachi Garella for driving all the way from Houston to record it.

Here’s a strict subset of the topics we covered:

  • The story of AI models solving the Navier-Stokes Millennium Problem, insofar as it’s known
  • Can recent AI proofs be called “truly creative”?
  • The history of AI before the LLM revolution
  • What do we mean when we call LLMs “black boxes”?
  • The achievements of the field of interpretability
  • What exactly happened in the OpenAI/HuggingFace incident
  • Must we avoid all “anthropomorphizing language” when discussing the HuggingFace incident? (spoiler alert: no)
  • Examples of major open problems in quantum computing theory that I cared about for decades and that AI models have recently solved
  • Effects of the current AI cataclysm on the math community, especially students
  • What annoys me the most when I listen to AI talks
  • My experiences at OpenAI, why they hired me, and the watermarking work that I did there

Enjoy!

More AI-related content coming soon, as this blog—like much of the rest of the world—continues its transition to “all AI, all the time” (except still 100% written by an aging, deteriorating biological brain)


And for those who just can’t get enough of my rocking back and forth, using too many filler words, as I explain theoretical computer science! Here’s a second podcast, this one mainly on quantum computing, with Seb Agertoft, who I thank for doing it. Enjoy!

Doug Natelson — Annual Nobel speculation thread

 

It's that time of year yet again.  The physics Nobel will be announced next Tuesday, and the chemistry prize on Wednesday.  Who will it be this time?  Please speculate in the comments.  As is my annual futile tradition, I will put forward that the physics prize could be Aharonov and Berry for geometric phases in physics (even though Pancharatnam is intellectually in there and died in 1969).  This is a long shot, as always, especially since last year's prize for “for the discovery of macroscopic quantum mechanical tunnelling and energy quantisation in an electric circuit” was condensed matter related.  Perhaps optical metamaterials, with Yablonovitch, Smith, and Pendry?  (Though that's another one where picking just three names is challenging.)  Astro-related is likely due, so perhaps something CMBR related?

Proofs and Prompts — Can AI truly prove anything?

Stepan Nesterov, PhD student at Stanford University

About a hundred years ago, David Hilbert gave a rigorous definition of a mathematical proof:

Definition. A mathematical proof is a sequence of formal statements A1,…,AnA_1, \ldots, A_n such that every statement AiA_i in the sequence either belongs to a predefined set of axioms of logic or there exists a statement AjA_j such that both AjA_j and Aj⇒AiA_j \Rightarrow A_i occur earlier in a sequence. A statement is provable, if there exists a mathematical proof whose last entry is that statement.

I will not recall here the definition of the set of axioms of logic, which was, of course, the main technical contribution here. At the same time, Church and Turing gave a rigorous definition of a computable function:

Definition. A function f:ℕ→ℕf \colon \mathbb{N} \to \mathbb{N} is called computable, if there exists a Turing machine MM which, when run on an otherwise empty tape with a binary expansion of a natural number nn printed, halts with an otherwise empty tape with a binary expansion of the natural number f(n)f(n) printed.

Can AI truly prove anything?

Let us first talk about the relationship between the informal notion of an algorithm, and the formal definition of computability. The philosophical statement that the two notions are equivalent, is known as the Church-Turing thesis. I am not aware of any counterexamples to this thesis being seriously proposed in the literature. Some authors have certainly entertained the possibility of a computing device falling into a black hole finishing the nn’th step of the computation in 2−n2^{-n} seconds, thereby completing an infinite computation in finite time. However, I don’t think that anybody believes in the physical feasibility of such computations: they are merely thought experiments designed to illustrate the blind spots in current physical theories.

The practicality of converting from informal algorithms to Turing machines, and vice versa, is an entirely different question. John von Neumann, while designing ENIAC, assumed that programmers would be low-skilled workers, because the job of converting informal instructions into machine code is purely mechanical. However, von Neumann’s assumption could not be further from the truth. The necessity of reading and modifying programs written by others led to the creation of assembly languages, and eventually high level programming languages. The desire to have a computer run multiple programs consecutively without the programmers coordinating in advance which sections of the memory they were allowed to use led to the invention of operating systems. As a result, by the end of the century, a typical Silicon Valley company had to employ many highly skilled and highly specialized workers in order to create a working software product. Even if the algorithms underlying something like an Uber app are mathematically trivial to describe, the creation of such an app needed a very large amount of man-hours.

On the other hand, the current revolution of artificial intelligence is directly based on creating programs which embody algorithms nobody can understand. Consider the following not quite mathematical questions:

  • Does there exist a computable strategy in the game of chess which can beat a human grandmaster?
  • Does there exist a computable function which outputs 1 on any picture of a cat and 0 on any picture of a dog?

Nowadays, anybody can use their computer to train a neural network which successfully accomplishes tasks such as these. The resulting program applies a large sequence of matrix multiplications and softmaxes to a digital representation of a picture or a chess position and returns a numerical answer. However, no human can look at these matrices and explain what actual chess insight the neural network uses to beat a human, or what exactly about the curvature of the lines and the color contrast of a picture distinguishes a cat from a dog.

If you look at how people talk about software nowadays, they remain firmly on the ‘formal’ side, not the ‘human understandable’ side. A program is understood as a tangible object which can run on your computer. There still remain some very serious concerns about using programs created by AI agents. If an AI piloting a car leads to a person’s death as a result of an accident, do the creators of the AI bear responsibility? Whatever the reason might be, these legal issues remain largely unresolved, but I’m sure that at some point they will be addressed.

What is a proof?

After this crash course on the history of computers, let us return, finally, to mathematics. Hilbert’s definition of a mathematical proof quickly became the de facto accepted definition of what a proof is. This gave mathematicians a possibility to brag about how mathematics is the only field of human knowledge which deals in completely objective truth. Unlike natural sciences, where the truth depends on believing that somebody did an experiment correctly, and therefore, ultimately, on the consensus of the scientific community, the mathematical proof is absolute.

Of course, all of this was always merely a convenient lie. To my knowledge, no mathematician ever attempted to submit an article with proofs strictly adhering to Hilbert’s system. In reality, most mathematicians will know that the standard axioms for set theory have the acronym ZFC, that Zorn’s lemma is equivalent to the axiom of choice but practically nothing else. One can simply look at the countless statements which changed status over the course of mathematical history to confirm that community consensus is very much a factor in deciding what constitutes a proof:

  • The XVII century had seen serious controversy over the fundamental theorem of algebra. Gauss’ thesis of 1799 criticized proofs by d’Alembert and Euler as not rigorous and relying on unproved assumptions. We now know, armed with a rigorous construction of ℝ\mathbb{R}, that these proofs were correct.
  • The Jordan curve theorem carries the name of Camille Jordan, who gave a proof of it in his analysis textbook in 1887. In the XX century, the consensus was that Jordan’s proof was flawed, and the first correct proof is due to Veblen in 1905. Then, in 2007, Hales has written an article defending Jordan’s proof, claiming that the only missing step was proving the theorem for polygons, which is easy. 
  • Many fundamental theorems in algebraic geometry carry the names of Enriques, Castelnuovo and Severi. They have published proofs which were at the time accepted in their community, even though the rigorous definitions of some of the basic concepts such as ‘generic point’ were absent from their work. When the rigorous definitions were later supplied by the work of van der Waerden, Noether, Zariski, Weil, many of the Italians’ original proofs became justified.
  • Elie Cartan’s original work on differential forms was not rigorous, but when de Rham has given a definition of a differential form, he discovered that all the identities that Cartan gave were correct. Yet people have cited Cartan’s work even before that point in time.

Conversely, when Appel and Haken proved the four-color theorem using computer assistance in 1976, their proof was controversial, because a human could not comprehend this proof in its entirety. Hales’ computer assisted proof of the Kepler conjecture was in review in Annals of Mathematics for seven years. The journal has even considered adding a special editorial note saying that the computer-related parts of the proof could not be verified by the journal’s editors. How would one explain this controversy, if one seriously subscribes to the Hilbert’s notion of a proof? Why would a computer have a problem generating a formal proof considering thousands of cases required? Indeed, both of these controversial proofs eventually received formalizations in Rocq, forcing mathematicians to begrudgingly accept that yes, both the four-color theorems and the Kepler conjecture have a formal proof in the sense of Hilbert, even though no human could understand them.

The classification of finite simple groups presents a different case study. While the proof is not computer assisted, it is very long, combining articles by many different people spanning in total thousands upon thousands of pages. Some mathematicians chose to reject the classification on the grounds that while every part of it was peer reviewed, no single person could possibly understand the proof in its entirety. There is currently no project dedicated to the formalization of the classification, but I’m sure that many mathematicians would begrudgingly accept its correctness if some AI agent produces suitable Lean code tomorrow. Again, from the point of view of Hilbert’s definition, this is very strange. He gave no requirement that the length of a proof should be somehow bounded by a thousand pages. In view of all that, I propose to consider taking the following alternative more seriously:

Definition. A mathematical proof is a text in natural language which explains to a person sufficiently well-trained in mathematics why a certain statement is correct. A statement is mathematically proven, if there is a mathematical proof which is accepted to be correct by sufficiently many well-respected members of the mathematical community.

The elephant in the room

Why am I saying all this? It is because our ego, our frivolous belief in the absoluteness of mathematics is currently being exploited by AI companies in the most cruel way. To them, mathematics is a problem to be solved, not unlike the game of chess. They hold an opinion that if an artificial agent could be created, which is significantly better than humans in producing Hilbert-style proofs, typically written today in the form of Lean code, then surely mathematics becomes a solved endeavor, and mathematicians will become obsolete. But in reality, the story’s not over when a formal proof of a statement is produced by any means necessary. When early programmers wrote pure machine code, they discovered that sometimes other people need to understand their programs in order to read and modify them. Then why, I ask, is it so difficult for us, mathematicians, to articulate why we need to understand the formal proof artifacts in order to practice our art?

To elaborate on my more realistic definition of a proof, I will discuss the following corollary:

Corollary. The statement ‘there exists a complex structure on S6S^6 was not mathematically proven on August 23, 2026.

As it stands, the situation is the following: on August 23, 2026, Claude Fable was prompted to construct a complex structure on a six-dimensional sphere. For reasons I don’t entirely understand, Mr. Fable does not announce his results by himself. Instead, he relies on his flesh puppet Levent Alpöge to announce the results on the personal X account in a maddeningly informal style. As an example of Alpöge’s sense of humor, one can refer to his announcement of the discovery of the Hadamard matrix of order 668. The announcement simply read

+++-+-+-+-+—+—+++-++——++—+—-+—–+++—++-+—+++++++-+–+-+-+–++-+–+–+—-+-+-+-+-++-++—+–++++–++-++++-+++++–+—++-+—…\ldots

and so on with no accompanying text whatsoever. In this case, Alpöge’s announcement was more human-like:

Please welcome to the world a beautiful new geometric object, to do with a problem i’ve always loved. claude really contains multitudes:D Does S6S^6 admit a complex structure? Yup

Alpöge then asked Fable’s less powerful sibling, Opus, to produce a description of Fable’s construction, which resulted in a 108-page pdf file. The very first page of that file plainly describes a mistake in the 2020 article by Campana, Demailly and Peternell. These authors have claimed to have a proof that the algebraic dimension of a complex manifold homeomorphic to S6S^6 must be zero, whereas Fable’s example has algebraic dimension one. Of course, the whole problem is famous to have inspired countless wrong solutions, so it’s entirely plausible that Campana, Demailly and Peternell has made a mistake. It is also plausible that Fable has made a mistake. It is worth noting that Alpöge himself is a number theorist, not a topologist. It seems that Fable’s construction is not so involved that a number theorist like Alpöge would be unqualified to evaluate it by himself. Yet, it is clear from the context of the announcement that he chose not to engage with the construction seriously, and delegated the work of digesting it to Claude Opus. If the construction is correct, then we have a problem of attribution, not unlike the problem of criminal responsibility for a car crash involving autopilot.

As trivial as the action of prompting Claude is, without Alpöge doing it, the proof wouldn’t have existed. Does Alpöge deserve credit as the discoverer of the complex structure on S6S^6? Qiaochu Yuan claims to have suggested this problem to Alpöge three days before the announcement. Does he deserve some credit as well? What about the team that worked on the development on Claude Fable?

Dealing with artificial mathematics

Because this question infringes on the domain of morality, I do not intend to answer it by supplying hard truths as arguments. Instead, I propose to consider what answer would we like to be accepted in order to benefit society as much as possible? Many mathematicians now wonder whether a day will soon come when an AI will solve the Riemann Hypothesis. So let me indulge in a little thought experiment and try to imagine what the mathematical community’s most likely response will be.

There is no way of knowing whether this hypothetical AI proof will rely on difficult computations, creating new, never before seen theories, or a combination of both. Given that the Riemann Hypothesis have resisted all attempts at proof for over 150 years, the AI proof will almost certainly be very long and difficult. Otherwise, somebody with a human brain would have probably found it already. Let’s say, for the sake of the argument, that AI proof runs for a 1000 pages. Immediately after the announcement, a large conference is organized. All the leading experts in analytic number theory join forces to try and make sense of the gigantic proof. After weeks of hard work, they are able to agree on what a few main insights of the proof are. In a year or two, the Proceedings from the conference are published, condensing the proof to about 500 pages and highlighting its main ideas in a human-readable way.

Now, these mathematicians interpreted the proof because of pure curiosity. In fact, it’s entirely possible that some of them had to compromise their job responsibilities because of it, say, by having to arrange that somebody substitute for the lecture they had to teach on the week of the conference. But the question then becomes: do we need to encourage AI developers to make more proofs, mathematicians to digest them, or a combination of both?

I believe that there are practical reasons to start making the following distinction: when a mathematical statement is known to have a formal proof because an AI system has generated it in Lean, we say that the statement is merely correct. We only say that a statement is mathematically proven, if there is a natural language proof which has been understood at the very least by the field’s leading experts to their satisfaction. I also believe that we should heavily encourage the work which leads to more statements being mathematically proven, and we should discourage the work which only leads to expanding the set of statements known to be merely correct.

My primary justification for this belief is the following: when a future AI system will write the proof of the Riemann Hypothesis, it will inevitably use the language of modern analytic number theory to do so. It would only be possible to understand this proof if you have previously studied analytic number theory, by, for example, reading textbooks written by humans. In a world where the emphasis is on the statements which are merely correct, eventually a point of no return will come. The proof of difficult theorems given by AI would only be able to be understood by another AI. It wouldn’t be impossible to verify that the AI is telling the truth by using Lean as a gold standard of truth in addition to AI cross-examination to obtain high-level descriptions of new mathematical ideas. But if no professional mathematicians are trained in universities, higher mathematics as a language would become extinct. 

The thing is, the language itself is just as valuable, if not more valuable, than knowing that such and such theorems is correct. In fact, the current AI revolution itself only became possible because the researchers in AI labs were trained, among other things, in advanced mathematics. If we were to blindly trust AI in proving our mathematical theorems, how is this meaningfully different from plainly saying that our ideal vision of the future is not ‘Star Trek’, but ‘Idiocracy’?

Conclusion

Fervent AI defenders would of course say that if we have a superintelligent system, then it will have no problem of explaining its genius solutions to the Millennium Problems to us lowly humans, therefore very quickly making them mathematically proven and not just merely correct. First of all, that is not the world we live in. OpenAI feels no obligation to train their most powerful models with a human-readable chain-of-thought, and Levent Alpöge feels no obligation to share Fable’s complete output as it solves the Jacobian conjecture, leaving us completely in the dark. But even if we were to live in that world, I do not believe that delegating all novel problem-solving to AI is a sustainable way of doing mathematics. As a historical comparison, we can look back to the time when calculators were invented. Our society eventually came to the conclusion that kids in elementary school need to learn to perform arithmetic operations by hand and familiarize themselves with its properties. In high school, they are then allowed to shortcut them using calculators, for they are well-equipped to perform sanity checks should they press a wrong button. Would it have mattered if somebody made a calculator with an enormous screen, which gives a complete long division as an output instead of just an answer? 

Currently there are few people that seriously attempt to reconcile human understanding with the current AI climate. Terrence Tao has multiple blogposts dedicated to digestions of whatever result OpenAI or Anthropic decided to work on this week. I believe that he is as good of a digester as he is only because he was already highly skilled in coming up with his own mathematical ideas. It truly does matter that enough mathematics is produced by using natural intelligence, or we risk losing the skill to do so.

Time will tell what the corresponding place of AI in mathematical education and research will be. But one thing is clear to me: prompting GPT Astra to manufacture unreadable Lean proofs for yet another Erdos problem has nothing to do with mathematics whatsoever. If you truly care about mathematics, make sure that mathematics is treated with the respect it deserves. Do not discuss AI-generated texts as if they were mathematical proofs. If you do discuss them, make sure to indicate whether a theorem is currently mathematically proved, merely correct, or, as is sometimes the case, claimed to be “proven” in a completely unverified AI slop file. Make sure that the button mashers from frontier AI labs, who only know the letter ‘X’, and not the letters ‘a’, ‘r’, ‘i’, ‘v’, do not dominate the mathematical discourse. Finally, learn and teach mathematics to make it last through the turbulent years of AI and the centuries beyond.


Received 5 September 2026.

September 28, 2026

Terence Tao — Joint Statement about Mathathon

[This is a guest post by a coalition of Caltech mathematicians: the original organizers of Mathathon (Alvan Arulandu, Andrea Li, Avni Garg, Brian Zhao, Caiman Moreno-Earle, Sathvik Redrouthu), two coauthors of the Open Letter about the Mathathon (Dylan King, Tasmin Chu), and other members of the Caltech community (Mayla Ward, Merrick Hua, Robert Joseph George, Sergei Gukov, Shaowu Zhang, Tony Yue Yu, Vihaan Dheer). This blog post was initially written in a different file format and converted using AI. It is also crossposted at Proofs and Prompts— T.]

Mathathon is centered around one question: how can we use AI to augment human understanding of mathematics?

This past week, the original organizers of Mathathon, two coauthors of an open letter about the mathathon, and other members of the Caltech community engaged in a conversation to redesign this event. We quickly reached a consensus: problem solving is but one component of mathematics; mathematical understanding and exposition are similarly meaningful. But these components have often been overlooked by grantmakers and hiring committees. We want to celebrate human mathematicians who undertake this work.

Thus the new theme of Mathathon is Old Problems, New Proofs.

Consider the four color theorem, the ABC conjecture, or the Navier-Stokes problem. Many mathematicians find their proofs—or claimed proofs—difficult to understand or unsatisfying. We invite our participants to pick a problem with an unintuitive solution, learn as much as possible in 40 hours, and present their findings to their peers. Then, they will take two months to develop an alternative proof or exposition. Afterward, they’ll submit an explainer and a GitHub repository. The explainer can take any form: a paper, a blog post, a video, an interactive game, et cetera. The repository will store whatever was used to produce the explainer—LLM chat histories, code used to produce visualizations, and more—so the mathematics community can examine the ideas behind the finished product.

By shifting the focus from open problems, we want to encourage Mathathon participants to explore everything else we do to make sense of mathematics: generalizations, new notations and techniques, connections to other problems. These innovations can be more exciting than solving individual problems.

Mathathon treats AI as one of many tools available to mathematicians. We no longer receive sponsorships from developers of proprietary AI models. Participants are given a cash grant to purchase any tool they need, whether that be LLM credits, HPC access, or pen and paper. Proprietary models are allowed, but we encourage the use of open-source tools. Our goal is to empower each participant to make their own decision about the tools used at Mathathon and in their own work.

The new event will take place on November 13–15. It is co-organized with the Foundation for Science and AI Research (SAIR), which will provide compute, host workshops, and offer prizes to teams that use open-source models. SAIR is also administering travel grants for participants. Our lead donor is XTX Markets, an algorithmic trading firm and a major supporter of academic research, open-source infrastructure, and community-led initiatives in AI-for-math.

The organizers and advisors of Mathathon range from AI proponents to critics, from undergraduates to well-established mathematicians. Working with this group has been a learning experience for everyone involved. We hope that it will bring together a similarly diverse set of judges and participants.

If you want to deepen your understanding of math, we hope to see you at Mathathon. Our mission is for you to learn from and teach your peers.

If you want to have fun, we hope to see you at Mathathon. It’s still a student-led event, run with the same playfulness that first inspired it.

If you are curious about AI capabilities, we hope to see you at Mathathon. Finding alternative proofs to solved problems and presenting them in human-oriented ways is a novel challenge for LLMs.

If you are concerned about AI’s impact on mathematics, we hope to see you at Mathathon. It celebrates human-oriented expositions. Mathathon requires participants to fully disclose and justify LLM usage and asks them to consider whether smaller-scale open-source tools would suffice.

Finally, Mathathon isn’t a conclusive guide to AI in mathematics. The organizers hold different opinions about AI-assisted problem solving, open-source versus proprietary models, and more. But we all believe that mathematicians can use AI to advance human understanding. Mathathon is a first experiment. We want to pave the way for future initiatives that explore AI’s roles in learning, teaching, review, and research, while placing human understanding and community building at the forefront. Mathematics is undergoing its biggest change in decades, and we must work together to adapt.

Acknowledgements. We are grateful to Terence Tao and Andrew Wu for reviewing this statement and leaving comments, as well as to countless others who offered feedback about Mathathon. All of the views above are solely our own and do not reflect those of Caltech or any other organizations we belong to.

Terence Tao — 247A, Notes 1: Rearrangement-invariant spaces

Disclaimer: due to current events, I have not been able to devote as much time to lecture notes preparation as I would have liked, so I apologize in advance for the unpolished nature of the text below, which has been largely recycled from previous lecture notes I have written.

This is the first set of lecture notes for my graduate course 247A, “Fourier analysis”. The course name is rather general, but I will focus the course not on the Fourier transform per se, but on the closely related topic of real variable harmonic analysis, with a particular emphasis on Calderón–Zygmund theory, which underlies basic tools in PDE such as the theory of Sobolev spaces.

To avoid confusion at the outset, let us make the distinction between real-variable harmonic analysis and abstract harmonic analysis, which are only distantly related to each other despite the similar names. Abstract harmonic analysis, roughly speaking, is the extension of the classical theory of the Fourier transform to other domains, such as locally compact abelian (LCA) groups, non-abelian Lie groups, or symmetric spaces, and typically involves a blend of representation theory, group theory, and analysis. Real-variable harmonic analysis, by contrast, tends to work on classical domains, such as a Euclidean space {{\bf R}^d}, a torus {{\bf R}^d/{\bf Z}^d}, or a lattice {{\bf Z}^d}, although many of the techniques can extend to more general domains (e.g., to Riemannian manifolds). While the Fourier transform often plays a prominent role (in particular, by setting the stage for time-frequency analysis and enabling various decompositions or other transforms that involve frequency space or phase space in addition to physical space), real-variable harmonic analysis is often focused on estimating other transforms or expressions that often interact well with the Fourier transform, but need not explicitly invoke it. Examples include the Hilbert transform

\displaystyle  Hf(x) := p.v. \int_{\bf R} \frac{f(y)}{x-y}\ dy \ \ \ \ \ (1)

(where we have made the somewhat arbitrary decision to omit the normalizing constant {\frac{1}{\pi}}) or the Hardy-Littlewood maximal function

\displaystyle  Mf(x) := \sup_{r>0} \frac{1}{|B(x,r)|} \int_{B(x,r)} |f(y)|\ dy.

These operators do not at first glance seem related to the Fourier transform, but as we shall see in later notes, the maximal function can be used to control the size of the Hilbert transform, and the Hilbert transform is a Fourier multiplier, with

\displaystyle  \widehat{Hf}(\xi) = -i \, \text{sgn}(\xi) \hat{f}(\xi)

for any reasonable function {f}, where we use the normalization {\hat f(\xi) := \int_{\bf R} f(x) e^{-2\pi i x \xi}\ dx}.

A typical question in harmonic analysis is the following: let {f} be some function on a standard domain (such as Euclidean space), and let {T f} be an explicit transform of {f} (e.g., the Hilbert transform {Hf} or maximal function {Mf}). To what extent is the “size” of {Tf} controlled by the “size” of {f}? The value of such bounds often lies in the general nature of the input function {f}; some mild regularity or decay hypotheses might be imposed on {f}, but beyond that the function {f} is typically not required to have a very structured form (in particular, it need not be describable by any closed-form expression).

In many situations the transform {T} being studied is linear or sublinear, in which case the natural type of bound to ask is a linear bound

\displaystyle  \|Tf\|_Y \leq C \|f\|_X

for suitable function space norms {X,Y} (e.g., {L^p} norms), and {C} is some bound. Depending on the application, we may be interested in various levels of precision regarding the bound {C}:
  • (a) Optimal bounds, in which we seek the exact optimal value of {C} (i.e., the operator norm of {T}). For instance, the optimal {L^2({\bf R}) \rightarrow L^2({\bf R})} constant for the Hilbert transform {H} is exactly {\pi}, and the optimal {L^1({\bf R}) \rightarrow L^{1,\infty}({\bf R})} constant for {M} in one dimension turns out to be {\frac{11+\sqrt{61}}{12}} (a result of Melas).
  • (b) Bounds accurate up to absolute constants (or maybe constants that can depend on basic parameters such as the ambient dimension).
  • (c) Bounds in which we are willing to accept “logarithmic type losses” such as {\log R} or {R^{o(1)}} in auxiliary parameters, such as a scale parameter {R}.

All three regimes are interesting, but we will focus in this class on the regime (b), where we can “afford” to lose absolute constants in the bounds, but will work hard to avoid any logarithmic losses. In particular, significant effort will be devoted in this class to avoiding “logarithmic pileups of scales”, in which the contributions of different dyadic scales such as {[2^k, 2^{k+1}]} for {k=1,2,3,\dots} all potentially contribute an equal amount that “interfere constructively” to cause a logarithmic divergence. This can be unnecessarily conservative when one is in regime (c) (which is for instance the situation in modern topics such as restriction theory or the Kakeya conjecture); nevertheless, the general skills gained by trying to not lose even a logarithmic factor in the bounds are often valuable in these other types of analysis.

When dealing with linear or sublinear problems, it is natural to try to decompose the initial function {f} into various smaller components {f_\alpha} by some decomposition {f = \sum_\alpha f_\alpha}, so that the transformed function {Tf} can be controlled by more tractable expressions {Tf_\alpha} in various ways (e.g., via the triangle inequality, by Bessel type inequalities, or by the more modern technique of decoupling inequalities). In short, the subject tends to proceed by a divide and conquer philosophy: it is generally preferable to replace a simple-looking but hard-to-estimate expression with a large, messy-looking combination of expressions that are easier to estimate. As such, the aesthetics of the subject are almost the reverse of those in the more algebraic portions of mathematics, in which progress is often made by making the expressions involved look as simple and unified as possible.

One of the main themes in this classical type of harmonic analysis is the struggle to understand the effect of two phenomena in integrals or sums: singularity and oscillation. The Hilbert transform (1) is a quintessential example of a singular integral, which combines both features: the non-locally integrable nature of the kernel {\frac{1}{x-y}} provides the singularity, but the sign change from {x < y} to {x>y} provides the oscillation. Classically, the interplay between these two phenomena can be tamed by analyzing the behavior of this operator both in the time (or “physical”) domain and in frequency (or “Fourier”) domain; in particular, the fact that singular integral operators such as the Hilbert transform are simultaneously a well-behaved Fourier multiplier and is “pseudo-local” in physical space lies at the heart of the standard Calderón–Zygmund theory for such operators. This dovetails nicely with more modern “time-frequency analysis” approaches to the subject, which can also handle other interesting operators, such as restriction or Bochner–Riesz operators via tools such as the wave packet decomposition, although these will be outside the scope of this course.

In this initial set of notes I will ignore the effect of oscillation, and develop some tools, such as interpolation theory, which can help control non-oscillatory sums and integrals if they are not too singular. Here, the focus will be on rearrangement-invariant spaces, such as the Lebesgue spaces {L^p} and their weak variants {L^{p,\infty}}, which are function spaces that are useful for measuring how “singular” or “decaying” various functions are, but do not pay attention to how they oscillate or where their mass is distributed. As such, these spaces do not capture the underlying geometry of the domain, which also plays an essential role in the subject; but it is nevertheless essential to have a good base understanding of the rearrangement-invariant theory before moving on to the more delicate aspects of harmonic analysis that are sensitive to rearrangements.

We will use the following asymptotic notation throughout the course: {X = O(Y)}, {X \lesssim Y}, or {Y \gtrsim X} denotes the assertion that {|X| \leq CY} for some constant {C}, and write {X \sim Y} for {X \lesssim Y \lesssim X}. If we permit this constant {C} to depend on some ambient parameters, we indicate this by subscripts; for instance, {X = O_{p,d}(Y)} or {X \lesssim_{p,d} Y} denotes a bound of the form {|X| \leq C_{p,d} Y} for some constant {C_{p,d}} that can depend on {p} and {d}. As indicated above, in this course we will generally not dwell much on exactly what these constants {C_{p,d}} are, or attempt to optimize them.

— 1. {L^p} norms —

Suppose one has some measurable function {f : X \rightarrow {\bf C}} on some measure space {(X,{\mathcal B},\mu) = (X,\mu)}. (Here we will follow the common practice if identifying functions that agree almost everywhere; in particular, we will be content to work with functions that are undefined on a set of measure zero. Also, while we work here with complex-valued functions throughout, most of the discussion here is also valid for real-valued or vector-valued functions.) Informally speaking, to measure how “big” such a function is, there are two (imprecisely defined) basic statistics to be aware of:

  • The height or amplitude of the function, which describes what the typical size of the magnitude {|f(x)|} is for {x} in the “dominant” component of the support of {f}; and
  • The width of the function, which describes the measure of this dominant component.

Example 1 (Informal) Given a Gaussian wave packet type function

\displaystyle  f(x) = A \exp( - |x-x_0|^2 / R^2 ) e^{2\pi i \xi \cdot x}

on {{\bf R}^d} for some {A, R > 0} and {\xi \in {\bf R}^d}, this function has magnitude {\approx A} on the ball {B(x_0,R)}, which has volume {\approx R^d} (if we allow constants in the informal {\approx} notation to depend on the dimension {d}), so such a function has height {\approx A} and width {\approx R^d}.

Example 2 (Informal) The function {\frac{1}{x}} on {{\bf R}}, which is implicitly involved in the definition of the Hilbert transform (1), does not have a clear amplitude or width as is. However, if one performs a dyadic decomposition

\displaystyle  \frac{1}{x} = \sum_{n \in {\bf Z}} 1_{2^n \leq |x| < 2^{n+1}} \frac{1}{x},

where we use {1_E} to denote the indicator of a statement {E} (equal to {1} when {E} is true and {0} otherwise), then each component

\displaystyle  1_{2^n \leq |x| < 2^{n+1}} \frac{1}{x}

of this decomposition has height {\approx 2^{-n}} and width {\approx 2^n}. Thus, while this function can be viewed as a superposition of components of various heights and widths, rather than a single such component.

These informal concepts of height and width are too imprecise to work with in practice. Experience has shown that a convenient proxy for these concepts are the {L^p(X,{\mathcal B}, \mu)} norms of a function {f}, defined for {0 < p < \infty} as

\displaystyle  \|f\|_{L^p(X,{\mathcal B},\mu)} := \Big( \int_{X} |f(x)|^p \, d\mu(x) \Big)^{1/p}

and for {p = \infty} as

\displaystyle  \|f\|_{L^\infty(X,{\mathcal B},\mu)} := \mathrm{ess\,sup}_{x \in X} |f(x)|

where {\mathrm{ess\,sup}} denotes the essential supremum of the function {x \mapsto |f(x)|} with respect to the measure {\mu}. Often we abbreviate {L^p(X,{\mathcal B},\mu)} as {L^p(X,\mu)}, {L^p(X)}, {L^p(\mu)}, or just {L^p} (and abbreviate {\|f\|_{L^p(X,{\mathcal B},\mu)}} as {\|f\|_p}) when the missing arguments are clear from context. (For instance, when working with Euclidean spaces {{\bf R}^d}, the measure {\mu} is understood to be Lebesgue measure, and {{\mathcal B}} the Lebesgue {\sigma}-algebra, unless otherwise specified.) In terms of the width {W} and height {H} of a function {f}, one heuristically has

\displaystyle  \|f\|_{L^p(X,\mu)} \approx H W^{1/p}

for both finite and infinite values of {p}, with the convention that {W^{1/\infty}} is equal to {1} when {W} is positive and {0} when {W} is zero. In the case of a step function {f = A 1_E} (where now {1_E(x) = 1_{x \in E}} is the indicator function of a measurable set {E}), this heuristic becomes exact:

\displaystyle  \|f\|_{L^p(X,\mu)} = A \mu(E)^{1/p}.

The function space {L^p(X,\mu)} is defined as the set of all measurable functions {f: X \rightarrow {\bf C}} for which the {L^p} norm is finite, up to almost everywhere equivalence, though we will often abuse notation by identifying a function with its almost everywhere equivalence class.

In the case where {X} is discrete and {\mu} is counting measure, we abbreviate {L^p(X,\mu)} as {\ell^p(X)}, or even just {\ell^p}.

Example 3 Let {\alpha > 0}. On a Euclidean space {{\bf R}^d}, the function {f(x) := |x|^{-\alpha} 1_{|x| > 1}} lies in {L^p({\bf R}^d)} (with a norm of {O_{p,\alpha,d}(1)}) if and only if {\alpha > d/p}, while the function {g(x) := |x|^{-\alpha} 1_{|x| \leq 1}} lies in {L^p({\bf R}^d)} (with a norm of {O_{p,\alpha,d}(1)}) if and only if {\alpha < d/p}. The function {|x|^{-\alpha}} does not lie in any {L^p({\bf R}^d)}, although it only fails “logarithmically” to lie in {L^{d/\alpha}}. Thus we see that control in {L^p} for high {p} rules out severe local singularities at a point, while control in {L^p} for low {p} rules out insufficiently rapid decay at infinity.

As is well known (see these previous notes) these function spaces enjoy many useful properties:

Theorem 4 (Basic properties of {L^p} spaces)
  • (i) The space {L^p(X,\mu)} is a Banach space when {1 \leq p \leq \infty}, a Hilbert space when {p=2}, and a topological vector space when {0 < p < 1}.
  • (ii) If {0 < p, q, r \leq \infty} obeys the scaling condition

    \displaystyle  \frac{1}{p} + \frac{1}{q} = \frac{1}{r} \ \ \ \ \ (2)

    then one has the Hölder inequality

    \displaystyle  \| fg \|_{L^r(X,\mu)} \leq \|f\|_{L^p(X,\mu)} \|g\|_{L^q(X,\mu)} \ \ \ \ \ (3)

    for any measurable {f,g: X \rightarrow {\bf C}} (here we adopt the usual conventions {0 \cdot +\infty = +\infty \cdot 0 = 0}). In particular, if {f \in L^p(X,\mu)} and {g \in L^q(X,\mu)}, then {fg \in L^r(X,\mu)}.
  • (iii) If {f \in L^p(X,\mu)} for some {1 \leq p \leq \infty}, then one has the duality relationship

    \displaystyle  \begin{array}{c} \|f\|_{L^p(X,\mu)} = \sup \left\{ \left| \int_X f(x) \overline{g(x)} \, d\mu(x) \right| : \right. \\ \left. g \in L^q(X,\mu), \|g\|_{L^q(X,\mu)} \leq 1 \right\}. \end{array} \ \ \ \ \ (4)

    where {q} is the conjugate exponent to {p}, defined by {\frac{1}{p} + \frac{1}{q} = 1}.

Remark 5 Closely related to (iii) is the fact that the dual of {L^p(X,\mu)} can be identified with {L^q(X,\mu)} when {1 \leq p < \infty} (with the additional hypothesis that {X} is {\sigma}-finite if {p=1}), but in practice the relation (4) will already be good enough for our purposes.

Remark 6 The Banach space property gives us the basic triangle inequality

\displaystyle  \| \sum_i f_i \|_{L^p(X,\mu)} \leq \sum_i \| f_i \|_{L^p(X,\mu)} \ \ \ \ \ (5)

for both finite and infinite collections of functions {f_i \in L^p(X,\mu)} when {1 \leq p \leq \infty}, where in the infinite case the assertion is that if the right-hand side is finite, then the series {\sum_i f_i} is absolutely convergent almost everywhere, and obeys the above inequality (so in particular is in {L^p(X,\mu)}). For {0 < p < 1}, this inequality fails (can you come up with a counterexample?), but one has the weaker {p}-triangle inequality

\displaystyle  \| \sum_i f_i \|_{L^p(X,\mu)}^p \leq \sum_i \| f_i \|_{L^p(X,\mu)}^p \ \ \ \ \ (6)

in this case, which follows easily from iterating the easy observation that {|z+w|^p \leq |z|^p + |w|^p} for any complex numbers {z,w}, which in turn ultimately stems from the complex triangle inequality {|z+w| \leq |z| + |w|} and concave nature of {t \mapsto t^p} for {t \geq 0}. In particular, for a finite sum {\sum_{i=1}^N f_i}, another application of Hölder’s inequality gives the quasi-triangle inequality

\displaystyle  \| \sum_{i=1}^N f_i \|_{L^p(X,\mu)} \leq N^{\frac{1}{p}-1} \sum_{i=1}^N \| f_i \|_{L^p(X,\mu)} \ \ \ \ \ (7)

for {0 < p < 1}, which is not too much worse than (5) when {N} is not too large.

Exercise 7 Give an example to show that the quantity {N^{\frac{1}{p}-1}} in (7) cannot be replaced by any smaller quantity.

Exercise 8 For {f} a simple function, verify that {\lim_{p \rightarrow \infty} \|f\|_p = \|f\|_\infty}, and that {\lim_{p \rightarrow 0} \|f\|_p^p = \mu( \mathrm{supp}(f) )}, where {\mathrm{supp}(f) := \{ x: f(x) \neq 0 \}}. For this reason, the measure of the support of {f} is sometimes referred to as the {L^0} norm of {f}, though it would be more accurate (though confusing) to refer to it as the {0^{th}} power of the {L^0} norm.

Remark 9 Note that Hölder’s inequality is not just symmetric under the homogeneities {f \mapsto cf} and {g \mapsto cg} of the functions, but also under the homogeneity {\mu \mapsto \lambda\mu} of the underlying measure. This latter symmetry demonstrates why the condition {\frac{1}{p}+\frac{1}{q}=\frac{1}{r}} is necessary. (The first two symmetries demonstrate why {f} appears the same number of times on both sides of the inequality, and similarly for {g}.)

In the case of Euclidean space, the measure homogeneity symmetry {\mu \mapsto \lambda \mu} is equivalent to the scaling symmetry {x \mapsto ax} for {a > 0}, as the Jacobian of this map is {\lambda = a^d}. But the point is that by manipulating the measure directly, one still enjoys this symmetry even when no scaling operation is present.

It is instructive to try to understand inequalities such as (3) using the height-width heuristic introduced previously. Suppose informally that {f,g,fg} have heights {H_f}, {H_g}, {H_{fg}} and widths {W_f}, {W_g}, {W_{fg}} respectively. Then one expects the heights to be related by the formula

\displaystyle  H_{fg} \approx H_f H_g.

What about the widths? Heuristically, the region that {fg} concentrates in ought to be a subset of the region that {f} concentrates in, so

\displaystyle  W_{fg} \lessapprox W_f

and similarly with {W_f} replaced by {W_g}. We can combine these bounds as

\displaystyle  W_{fg} \lessapprox \min(W_f, W_g).

The bound (3) then is morally

\displaystyle  H_{fg} W^{1/r}_{fg} \lessapprox H_f W^{1/p}_{f} \cdot H_g W^{1/q}_g,

which on applying the previous bounds and (2) should simplify to

\displaystyle  \min(W_f, W_g)^{1/p} \min(W_f, W_g)^{1/q} \lessapprox W^{1/p}_{f} W^{1/q}_g.

But this is clear by bounding {\min(W_f,W_g)} by {W_f} for the first factor on the left-hand side, and by {W_g} for the second factor. Thus we see that the key geometric input that is morally driving the Hölder inequality is the simple fact that the concentration region of the product is contained in the concentration regions of the factors.

Exercise 10 Determine the cases for which (3) holds with equality (dealing with edge cases such as when one or more of {p,q,r} equal infinity as appropriate). Discuss how your conclusions align with the heuristic analysis presented above.

Exercise 11 If {1 < p < \infty}, determine the cases for which (5) holds with equality. What changes when {p=1} or {p=\infty}?

Exercise 12 Show that Hölder’s inequality is equivalent to the log-convexity of {L^p} norms:

\displaystyle  \begin{array}{c} \|f\|_r \leq \|f\|^{1-\theta}_p \|f\|_q^\theta \hbox{ whenever } 0 < p < q < \infty, 0 < \theta < 1 \\ \hbox{ and } \frac{1}{r} = \frac{1-\theta}{p} + \frac{\theta}{q}. \end{array} \ \ \ \ \ (8)

(For technical reasons one needs to first reduce to the case where {X} has finite measure, and then {f} and {g} are everywhere non-vanishing simple functions. Now consider the convexity of {|f|^\alpha |g|^\beta} with respect to a measure {|f|^\gamma |g|^\delta \mu} for some suitable exponents {\alpha,\beta,\gamma,\delta}.)

Exercise 13 (Direct approach to log convexity) Differentiate {\log \|f\|_p} twice with respect to {\alpha := 1/p} and show that this is non-negative (take {f} to be a non-zero simple function with finite measure support to avoid technicalities). This is an example of a monotonicity formula method — deriving estimates from a monotonicity property, which in turn follows from the non-negativity of a derivative.

You will see that this approach is surprisingly messy. For a much slicker proof, observe that (8) enjoys homogeneity symmetry in both {f} and {\mu}, which lets one normalise both {\|f\|_p} and {\|f\|_q} to equal one. Thus the task is now to show that if {\|f\|_p=\|f\|_q=1}, then {\|f\|_r \leq 1} for all {r} between {p} and {q}. This can be done by the pointwise convexity of {p \mapsto |f(x)|^p}, or more precisely the estimate

\displaystyle  |f(x)|^r \leq (1-\theta) |f(x)|^p + \theta |f(x)|^q;

the observant reader will note that this is merely the proof of Hölder’s inequality in disguise.

Let us now give a more unusual proof of the log-convexity which does not appeal to any pointwise convexity estimate, instead combining the “divide and conquer” strategy with an elegant (and rather cheeky) “tensor power trick“. Again normalise {\|f\|_p = \|f\|_q = 1}. We split {f} into a broad flat piece and a narrow tall piece

\displaystyle  f = f 1_{|f| \leq 1} + f 1_{|f| > 1}

which are disjoint, and thus

\displaystyle  \|f\|_r^r = \int_{|f| \leq 1} |f|^r + \int_{|f| \geq 1} |f|^r.

What we are doing here is exploiting some very basic intuition about {L^p} norms, namely that {L^p} bounds for large {p} tend to exclude tall narrow spikes, whereas {L^p} bounds for small {p} tend to exclude short broad tails. Of course, either sort of bound would exclude tall broad functions, and neither excludes narrow short functions. Once again, this intuition can be buttressed by considering the special case of step functions.

When {|f| \leq 1}, then {|f|^r \leq |f|^p}, and when {|f| \geq 1}, then {|f|^r \leq |f|^q}. Thus we end up with

\displaystyle  \|f\|_r^r \leq \int_X |f|^p + \int_X |f|^q = 2.

The above argument (which is a prototype of the real interpolation method) obtained an estimate which is off by a factor of two from what we wanted; this is a typical feature of the method. However we can recover this factor for free by the following tensor power trick. Let {M} be a large integer. We replace the measure space {(X,{\mathcal B},\mu)} by its {M^{th}} power {(X^M, {\mathcal B}^{\oplus M}, \mu^{\oplus M})} using the product measure construction, and similarly replace {f} with its tensor power {f^{\oplus M}: X^M \rightarrow {\bf C}}, defined by

\displaystyle  f^{\oplus M}(x_1,\ldots,x_M) := f(x_1) \ldots f(x_M).

One then observes that

\displaystyle  \begin{array}{rl}  \|f^{\oplus M}\|_{L^p(X^M)} &= \|f\|_{L^p(X)}^M = 1; \\ \|f^{\oplus M}\|_{L^q(X^M)} &= \|f\|_{L^q(X)}^M = 1; \\ \|f^{\oplus M}\|_{L^r(X^M)} &= \|f\|_{L^r(X)}^M. \end{array}

Now we apply the preceding arguments to {f^{\oplus M}} instead of {f} to deduce that

\displaystyle  \|f^{\oplus M}\|_{L^r(X^M)}^r \leq 2

which on taking {M^{th}} roots gives

\displaystyle  \|f\|_{L^r(X)}^r \leq 2^{1/M}.

Now the left-hand side is independent of {M}; take limits as {M \rightarrow \infty} and we obtain {\|f\|_r \leq 1} as desired.

The tensor power trick can be viewed as another application of symmetry: if an estimate is invariant under raising to a tensor power, then one can automatically replace all absolute constants with {1}; thus we obtain the “free lunch” of deducing a bound with an explicit constant {1}, from a bound with an unspecified constant (or even with “logarithmic losses”). Contrapositively, if an estimate is invariant under tensor power, then a weak counterexample (which shows that the constant must exceed one) can be amplified into a strong counterexample (which shows that no finite constant suffices) by tensor powering. The tensor power trick seems like a magical trick at present, but is actually exploiting some basic results in information theory such as the Shannon entropy inequalities and the central limit theorem; it also combines well with virtually any inequality which involves Gaussians. Unfortunately due to lack of time we will not be discussing these beautiful topics further in this course. At any rate one sees the power of abstraction in this tensor power trick. (One could similarly perform this trick in {{\bf R}^d}, so long as the constants only grew sub-exponentially in the dimension {d}.)

The final proof of log-convexity of the norm that we give here proceeds via complex analysis, and the maximum principle — which in many ways is a complex analogue of convexity (or subharmonicity). We need the following result from complex analysis, namely a form of the Phragmén–Lindelöf principle.

Lemma 14 (Three lines lemma) Let {f} be a complex-analytic function on the strip {\{ 0 \leq \Re(z) \leq 1 \}}, which is of at most double-exponential growth, or more precisely {|f(z)| \lesssim_f e^{O_f(e^{(\pi-\delta)|z|})}} for some {\delta > 0}. Suppose that we have the bounds {|f(z)| \leq A} when {\Re(z) = 0} and {|f(z)| \leq B} when {\Re(z) = 1}. Then we have {|f(z)| \leq A^{1-\Re(z)} B^{\Re(z)}} for all {z} in the strip.

Remark 15 The rather strange sub-double-exponential hypothesis here is completely sharp, as the example {f(z) = e^{-ie^{\pi iz}}} shows. Note in this hypothesis that we allow the implied constants in the asymptotic notation to depend on {f}, but the hypothesis is qualitative rather than quantitative: the value of these constants is irrelevant for the final conclusion, as long as they are finite. In practice, these sorts of qualitative hypothesis are usually easy to establish (especially when compared to quantitative estimates) by restricting, smoothing, or damping to a nice class of functions, or by smoothing out or discretising various operators and domains. See for instance the proof of this very lemma in which we upgrade “for free” a weak qualitative bound (sub-double-exponential growth) to a strong qualitative bound (decay at infinity).

Proof: The hypotheses and conclusion of the lemma are invariant under the operation of multiplying {f} by a constant (and adjusting {A,B} appropriately). So we may normalise {A=1}. Similarly, the hypotheses and conclusion of the lemma are invariant under the operation of multiplying {f} by an exponential {\exp(cz)} for some real {c}. Using this, one can also normalise {B=1}. So now {f} is bounded by {1} on both sides of the strip and we want to show it is bounded by {1} inside the strip.

Let us first assume that {f} is much better than exponential growth, namely that it goes to zero at infinity. Then for all sufficiently large rectangles {\{ 0 \leq \Re(z) \leq 1; -N \leq \Im(z) \leq N\}} the complex-analytic function {f} is bounded by {1} on all four sides of this rectangle, and hence in the interior also by the maximum principle, and we are done by setting {N \rightarrow \infty}.

Now let us handle the general case; as is usual when removing a qualitative assumption, we do this by a limiting argument. We replace {f(z)} by {f(z) \exp(\varepsilon e^{i [(\pi-\varepsilon)z + \varepsilon/2]})}; a little complex arithmetic shows that this converts the almost double-exponentially growing function {f} to one which is still complex analytic but is now decaying at infinity. It is still bounded by {1} at both sides of the strip, and hence by {1} in the interior also by the previous argument. Now take {\varepsilon \rightarrow 0} to conclude the claim. \Box

Exercise 16 Suppose that {f} is analytic on the strip {\{ 0 \leq \Re(z) \leq 1 \}}, obeys the sub-double-exponential bound {|f(z)| \lesssim_f e^{O_f(e^{(\pi-\delta)|z|})}} on the strip, and obeys the polynomial bounds {|f(z)| \lesssim (1 + |z|)^{O(1)}} on the sides of the strip. Show that it obeys the polynomial bound {|f(z)| \lesssim (1 + |z|)^{O(1)}} on the interior of the strip also.

To apply the three-lines lemma to prove (8), take {f} to be a simple function (with finite measure support) and consider the entire function

\displaystyle  z \mapsto \int_X |f|^{z}.

This function has exponential growth at most (because of the qualitative assumption that {f} is simple with finite measure support), and is bounded by {1} on the lines {\Re(z) = p} and {\Re(z) = q}, and hence (by a trivially rescaled version of the three lines lemma) bounded by {1} on the strip inside the lines. In particular it is bounded by {1} at {z=r}, which gives the claim for simple functions. The claim for more general functions (dropping the qualitative assumption of simpleness and finite measure support) then follows by a standard limiting argument (using for instance monotone convergence) which we leave as an exercise.

The above argument is a prototype of the complex interpolation method. As one can see, it can give slightly sharper results than the real interpolation method (though using the tensor power trick the real method can sometimes “catch up”), but on the other hand requires the quantities being studied to depend complex-analytically on a parameter rather than (say) real-analytically.

Having conclusively demonstrated the log-convexity (8) in multiple ways, let us now give some quick applications. It shows that control on two extreme {L^p} norms implies control of the intermediate {L^p} norms. Under additional assumptions on the measure space {(X, \mu)}, one of these extremes is not necessary. If the measure space is finite in the sense that {\mu(X) < \infty} (thus prohibiting functions from being arbitrarily broad), then higher {L^p} norms control lower ones:

\displaystyle  \| f \|_p \leq \|f\|_q \mu(X)^{\frac{1}{p} - \frac{1}{q}} \hbox{ whenever } 0 < p \leq q \leq \infty. \ \ \ \ \ (9)

Indeed this is trivial when {q=\infty}, and the general case then follows by convexity. The bound (9) can also be usefully written in terms of averages: if we write {\rlap{\hspace{2pt}-}\int_X f\ d\mu} for {\frac{1}{\mu(X)} \int_X f\ d\mu}, then we see that higher {L^p} averages control lower {L^p} averages:

\displaystyle  (\rlap{\hspace{2pt}-}\int_X |f|^p\ d\mu)^{1/p} \leq (\rlap{\hspace{2pt}-}\int_X |f|^q\ d\mu)^{1/q} \hbox{ whenever } 0 < p \leq q \leq \infty.

One way to view this is that as one lowers the exponent {p}, the exceptionally large values of {f} become less important, leaving the small values of {f} to dominate. By restricting {X} to its support {\mathrm{supp}(f)} one can refine (9) to

\displaystyle  \| f \|_p \leq \|f\|_q \mu(\mathrm{supp}(f))^{\frac{1}{p} - \frac{1}{q}} \hbox{ whenever } 0 < p \leq q \leq \infty.

(Note that this is a limiting case of log-convexity at the exponent {0}, in view of Exercise 8). In the converse direction, if the measure space is granular in the sense that one has a lower bound {\mu(E) \geq c} for all sets {E} of positive measure, then functions are prohibited from being arbitrarily narrow, and lower norms control higher norms:

\displaystyle  \| f \|_q \leq \|f\|_p c^{\frac{1}{q} - \frac{1}{p}} \hbox{ whenever } 0 < p \leq q \leq \infty.

This can be seen by first checking the {q=\infty} case, and then using log-convexity to get the remaining cases. In particular, in the {\ell^p} spaces we see that {\|f\|_{\ell^q} \leq \|f\|_{\ell^p}} for {q \geq p}. For {\ell^p} spaces on {N} points, we thus have (non-matching) upper and lower bounds

\displaystyle  \|f\|_{\ell^q} \leq \|f\|_{\ell^p} \leq N^{\frac{1}{p}-\frac{1}{q}} \|f\|_{\ell^q}, \ \ \ \ \ (10)

{\ell^p} and {\ell^q} norms are comparable to some extent, but the comparability gets worse as {N \rightarrow \infty} or as {p} and {q} get further apart.

Exercise 17 Heuristically justify the bounds (9), (10) by appealing to the informal notions of width and height of a function.

Exercise 18 When does equality occur for either of the inequalities in (10)? Note how the example that attains the lower bound is in many ways the “opposite extreme” to the example which attains the upper bound.

Lebesgue measure on Euclidean spaces {{\bf R}^d} with the usual Borel or Lebesgue {\sigma}-algebra is not granular. However one can create granularity by coarsening the {\sigma}-algebra. For instance, if we let {{\mathcal B}_1} be the {\sigma}-algebra generated by the lattice unit cubes {n + [0,1)^d} for {n \in {\bf Z}^d}, then we have granularity with constant {c=1}, and now lower {L^p} norms of functions {f} control higher ones — but only for functions which are measurable with respect to this algebra, i.e. only for functions which are constant on each lattice unit cube. (This is the first time we have actually manipulated the {\sigma}-algebra {{\mathcal B}} to say something non-trivial, as opposed to manipulating {f}, {X}, or {\mu}.) Thus we see that local constancy of functions can lead to additional estimates on {L^p} norms. Later on we shall see that frequency localisation achieves a similar effect as local constancy, as quantified by Bernstein’s inequality; this is a concrete manifestation of the famous Heisenberg uncertainty principle.

Finiteness and granularity of the measure space prevent a function from being too broad or too narrow respectively. Similar things happen when a function is being prevented from being too tall or too short; for instance if {f} is bounded above by a constant {M}, then we have

\displaystyle  \|f\|_q \leq \|f\|_p^{p/q} M^{1-p/q} \hbox{ whenever } 0 \leq p \leq q \leq \infty

(this is just log-convexity at the {\infty} exponent), while if {f} is bounded below by {M} on its support, then we have the reverse inequality

\displaystyle  \|f\|_p \leq \|f\|_q^{q/p} M^{1-q/p} \hbox{ whenever } 0 \leq p \leq q \leq \infty.

There are two obvious algebraic identities involving {L^p} norms which are worth knowing. The first is that one can interchange {l^p} sums with {L^p} integrals for any {0 < p < \infty}, in the sense that

\displaystyle  \| (\sum_n |f_n|^p)^{1/p} \|_{L^p} = (\sum_n \|f_n\|_{L^p}^p)^{1/p};

this is just an application of the Fubini-Tonelli theorem. Secondly, exponents can pass through {L^p} norms by changing the exponent: for any {0 < p,q < \infty} we have

\displaystyle  \| |f|^q \|_{L^p} = \| f \|_{L^{pq}}^q.

We shall use both of these identities in the sequel without further comment.

Exercise 19 Establish the bound

\displaystyle  \| (\sum_n |f_n|^q)^{1/q} \|_{L^p} \leq (\sum_n \|f_n\|_{L^p}^q)^{1/q}

for any measurable {f_n} and any {0 < q < p \leq \infty}.

— 2. Lorentz spaces —

Recall that the weak {L^p} norm {\|f\|_{L^{p,\infty}(X,\mu)}} of a function {f} is defined for {0 < p < \infty} as

\displaystyle  \|f\|_{L^{p,\infty}(X,\mu)} := \sup_{\lambda > 0} \lambda \mu( \{ |f| \geq \lambda \} )^{1/p}.

Since

\displaystyle  \|f\|_p^p = \int_X |f|^p\ d\mu \geq \int_X \lambda^p 1_{|f| \geq \lambda}\ d\mu = \lambda^p \mu( \{ |f| \geq \lambda \} )

for any {f} and {\lambda}, we obtain Chebyshev’s inequality

\displaystyle  \|f\|_{L^{p,\infty}(X,\mu)} \leq \|f\|_{L^{p}(X,\mu)}

(the case {p=1} is also known as Markov’s inequality). When {p=\infty} we adopt the convention that {L^{\infty,\infty} = L^\infty}.

We define weak {L^p(X,\mu)} or {L^{p,\infty}(X,\mu)} to be the space of all functions with finite {L^{p,\infty}(X,\mu)} norm, with the usual abbreviations. We sometimes refer to {L^p} as strong {L^p} to distinguish it from weak {L^p}.

Example 20 On a Euclidean space {{\bf R}^d}, the power function {|x|^{-\alpha}} lies in weak {L^p} if and only if {\alpha = d/p}. Indeed one can think of a weak {L^p} function as a function which is pointwise dominated in magnitude by a rearrangement of (a multiple of) {|x|^{-d/\alpha}}.

Suppose {0 < p < \infty}. From elementary calculus we have

\displaystyle  |f(x)|^p = p \int_0^\infty 1_{|f(x)| \geq \lambda} \lambda^p \frac{d\lambda}{\lambda}

and hence on integration and Fubini’s theorem

\displaystyle  \|f\|_p^p = p \int_0^\infty \mu( \{ |f| \geq \lambda \} ) \lambda^p \frac{d\lambda}{\lambda}.

To summarise, we have

\displaystyle  \|f\|_{L^{p,\infty}} = \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^\infty({\bf R}^+,\frac{d\lambda}{\lambda})}

and

\displaystyle  \|f\|_{L^p} = p^{1/p} \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^p({\bf R}^+,\frac{d\lambda}{\lambda})}

for {0 < p < \infty}. These two identities motivate introducing the Lorentz (quasi-)norm {L^{p,q}(X,\mu)} for {0 < p < \infty} and {0 < q \leq \infty} by

\displaystyle  \|f\|_{L^{p,q}(X,\mu)} := p^{1/q} \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^q({\bf R}^+,\frac{d\lambda}{\lambda})}. \ \ \ \ \ (11)

Thus for instance {L^p} norm is identical to the {L^{p,p}} norm. We shall abbreviate {\|f\|_{L^{p,q}(X,\mu)}} by {\|f\|_{L^{p,q}(X)}}, {\|f\|_{L^{p,q}}}, or even {\|f\|_{p,q}} when there is no chance of confusion.

Remark 21 For various reasons it is not worth trying to define Lorentz norms when {p=\infty}, although we will use the convention {L^{\infty,\infty}=L^\infty}. The most important values of {q}, in descending order, are {q=p}, {q=\infty}, {q=1}, and {q=2}; the other cases essentially never occur in applications.

Remark 22 The factor {p^{1/q}} is inconsequential, but is traditional in order to maintain compatibility with the strong {L^p} norm. But in practice the exact form of the Lorentz norm is not important; there are many formulations which are equivalent up to constants, and one generally just picks the formulation which is most convenient. The measure {\frac{d\lambda}{\lambda}} is of course multiplicative Haar measure on {{\bf R}^+}. One can interpret the equivalence of (i)-(iii) below by making the change of variables {\lambda = 2^m}, so that the Haar measure just becomes Lebesgue measure in {m} (modulo an inessential constant) and then passing from continuous {m} to discrete {m}.

Exercise 23 If {f: {\bf R}^+ \rightarrow {\bf R}^+} is a monotone non-increasing function, show that

\displaystyle  \|f\|_{L^{p,q}(X,\mu)} = \| f(t) t^{1/p} \|_{L^q({\bf R}^+, \frac{dt}{t})}.

(Depending on how you prove this, it may be convenient to first prove this for smoother {f}, such as diffeomorphisms with strictly negative derivative, in order to apply an inverse function theorem.)

Exercise 24 Show that a step function of height {H} and width {W} has an {L^{p,q}} norm of {(p/q)^{1/q} H W^{1/p}} for any {0 < p < \infty} and {0 < q \leq \infty}.

Exercise 25 For any {0 < p,r < \infty} and {0 < q \leq \infty} show that

\displaystyle  \| f^r \|_{L^{p,q}(X,\mu)} \sim_{p,q,r} \|f\|_{L^{pr,qr}(X,\mu)}^r.

It is obvious that these norms are both rearrangement-invariant and monotone. To get a better intuitive handle on what the {L^{p,q}} norm represents, we need some more definitions.

  • A sub-step function of height {H} and width {W} is any function {f} supported on a set {E} with the bounds {|f(x)| \leq H} almost everywhere and {\mu(E) \leq W}. (Thus {|f| \leq H 1_E}.)

  • A quasi-step function of height {H} and width {W} is any function {f} supported on a set {E} with the bounds {|f(x)| \sim H} almost everywhere on {E}, and {\mu(E) \sim W}. (Thus {|f| \sim H 1_E}.)

Remark 26 It is a little dangerous to put fuzzy notation such as {\sim} within a definition; if multiple quasi-step functions appear in an argument, the question then arises as to whether the implied constants are uniform. In our applications, the implied constants here are true absolute constants (like {1} and {2}) so this will not be an issue.

Remark 27 From the binary expansion of the unit interval {[0,1)} we see that a non-negative sub-step function {f} of height {1} and width {W} can always be decomposed as {\sum_{k=1}^\infty 2^{-k} f_k} where {f_k} is an actual step function of height {1} and width at most {W}. By homogeneity we have a similar statement for other heights. Because of this, bounds on step functions tend to automatically extend to sub-step functions (and hence quasi-step functions) without difficulty.

Just like actual step functions, the {L^{p,q}} norm of sub-step and quasi-step functions are well controlled; a sub-step function of height {H} and width {W} has {L^{p,q}} norm {O_{p,q}( H W^{1/p} )}, while a quasi-step function has {L^{p,q}} norm {\sim_{p,q} H W^{1/p}} – almost exactly like actual step functions. In the converse direction, it turns out that every {L^{p,q}} function can be decomposed as an {l^q} sum of “very different” {L^{p,q}}-normalised sub-step or quasi-step functions.

Theorem 28 (Characterisation of {L^{p,q}}) Let {f} be a function, let {0 < p < \infty} and {1 \leq q \leq \infty}, and let {0 < A < \infty}. Then the following are equivalent up to changes in the implied constants:

Remark 29 The formulations (ii), (iv) are useful when trying to use an {L^{p,q}} bound on {f}; the formulations (iii), (v) are useful when trying to obtain an {L^{p,q}} bound on {f}. Heuristically, the above theorem is trying to say the following. If {f} is a quasi-step function of height {H} and width {W}, then {\|f\|_{L^{p,q}} \sim_{p,q} H W^{1/p}}. But if {f} is instead the sum {\sum_n f_n} of quasi-step functions of height {H_n} and width {W_n}, and either the heights or the widths are sufficiently variable in {n} (e.g. one or the other grows like a power of two), then {\|\sum_n f_n \|_{L^{p,q}} \sim_{p,q} \| H_n W_n^{1/p} \|_{\ell^q_n}}.

Proof: We may use homogeneity symmetry to normalise {A=1}. The implications {(ii) \implies (iii)} and {(iv) \implies (v)} are trivial. To see that (i) implies (ii), set {f_m := f 1_{2^{m-1} < |f| \leq 2^m}} and {W_m := \mu( \{ 2^{m-1} < |f| \leq 2^m \} )}. (This is the “vertically dyadic layer cake decomposition”.) The only thing that requires nontrivial verification is (12); but one easily verifies that

\displaystyle  2^m W_m^{1/p} \lesssim_{p,q} \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^q([2^{m-2}, 2^{m-1}],\frac{d\lambda}{\lambda})}

and the claim follows by summing this in {l^q}.

Similarly, to see that (i) implies (iv), define

\displaystyle  H_n := \inf \{ \lambda: \mu( \{ |f| > \lambda \} ) \leq 2^{n-1} \};

note that this is a non-increasing function of {n}, which goes to zero as {n \rightarrow \infty} (this comes from the hypothesis that {\|f\|_{L^{p,q}}} is finite). We then define

\displaystyle  f_n := f 1_{ H_n \geq |f| > H_{n+1} }.

(This is the “horizontally dyadic layer cake decomposition”.) The only non-trivial thing to verify is (14). But one easily verifies the telescoping estimate

\displaystyle  \begin{array}{rl}  H_n 2^{n/p} &= (H_n^q 2^{nq/p})^{1/q} \\ &= (\sum_{k=0}^\infty (H_{n+k}^q - H_{n+k+1}^q) 2^{nq/p} )^{1/q} \\ &\lesssim_{p,q} (\sum_{k=0}^\infty 2^{-kq/p} \| \lambda 2^{(n+k)/p} \|_{L^q([H_{n+k+1}, H_{n+k}],\frac{d\lambda}{\lambda})}^q)^{1/q}\\ &\lesssim_{p,q} (\sum_{k=0}^\infty 2^{-kq/p} \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^q([H_{n+k+1}, H_{n+k}],\frac{d\lambda}{\lambda})}^q)^{1/q} \end{array}

and the claim follows by summing this in {\ell^q} and interchanging the summation signs. (We leave to the reader how to modify the above argument to handle the case {q=\infty}.)

It remains to show that (iii) implies (i) and (iv) implies (i). Suppose first that (iii) holds. It is clear that for any {m} we have

\displaystyle  \mu( \{ |f| > 2^m \} ) \leq \sum_{k=0}^\infty \mu(E_{m+k})

and hence

\displaystyle  \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^q((2^m,2^{m+1}],\frac{d\lambda}{\lambda})} \lesssim_{p,q} 2^m (\sum_{k=0}^\infty \mu(E_{m+k}))^{1/p}

and so on taking {l^q} summation in {m} it would suffice to show that

\displaystyle  \| 2^m (\sum_{k=0}^\infty \mu(E_{m+k}))^{1/p} \|_{\ell^q_m} \lesssim_{p,q} 1.

Raising to the {p^{th}} power we rewrite as

\displaystyle  \| \sum_{k=0}^\infty 2^{pm} \mu(E_{m+k}) \|_{\ell^{q/p}_m} \lesssim_{p,q} 1.

But from the hypothesis we have

\displaystyle  \| 2^{pm} \mu(E_{m}) \|_{\ell^{q/p}_m} \lesssim_{p,q} 1

and hence on shifting {m} by {k}

\displaystyle  \| 2^{pm} \mu(E_{m+k}) \|_{\ell^{q/p}_m} \lesssim_{p,q} 2^{-kp}.

The claim then follows from (5) or (7).

Now suppose that (iv) holds. Observe that for any {\lambda > 0} we have

\displaystyle  \mu( \{ |f| > \lambda \} ) \lesssim_{p,q} \sup \{ 2^n: H'_n \geq \lambda \}

where {H'_n} are the modified heights

\displaystyle  H'_n := \sum_{k=0}^\infty H_{n+k}.

Indeed, if {\mu( \{ |f| > \lambda \} ) > 2^{n-1}} for some {n} then one easily verifies that {H_n \geq \lambda} and hence {H'_n \geq \lambda}. The shifting trick and triangle inequality argument used previously shows that {H'_n} obeys the same bound (14) as {H_n}, thus

\displaystyle  \| H'_n 2^{n/p} \|_{\ell^q_n({\bf Z})} \lesssim_{p,q} 1.

We now compute

\displaystyle  \begin{array}{rl}  \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^q({\bf R}^+,\frac{d\lambda}{\lambda})}^q &\lesssim_{p,q} \int_0^\infty \lambda^{q-1} \sup \{ 2^{nq/p}: H'_n \geq \lambda \}\ d\lambda \\ &\lesssim_{p,q} \sum_n \int_0^\infty \lambda^{q-1} 2^{nq/p} 1_{H'_n \geq \lambda}\ d\lambda \\ &\sim_{p,q} \sum_n 2^{nq/p} (H'_n)^q \\ &\lesssim 1 \end{array}

as desired. (We leave to the reader how to modify the above argument to handle the case {q=\infty}.) \Box

Remark 30 For future reference we make the technical remark that if {f} is a simple function, then in the horizontal and vertical decompositions in the above theorem, only finitely many of the {f_n} are non-zero.

Remark 31 Suppose that the ratio between the tallest height and lowest non-zero height of a function {f} is {N} (i.e. there exists {A} such that {A \leq |f(x)| \leq AN} whenever {f(x)} is non-zero). Then the above theorem shows that two different Lorentz norms {\|f\|_{L^{p,q_1}}}, {\|f\|_{L^{p,q_2}}} with the same primary exponent {p} only differ by multiplicative powers of {\log N}. Similarly if the broadest width and narrowest width of a function differs by {N} (e.g. if {\mu(X)} is equal to {N} times the granularity {c} of {X}). What this indicates is that the secondary exponent {q} in the Lorentz norms only offers “logarithmic correction” to the Lebesgue norms {L^p}; in contrast, (10) shows that varying the primary exponent {p} leads to polynomial-strength changes in the norm. So as a first approximation (ignoring logarithms) one can pretend that {L^{p,q} \approx L^p}. Note also that for quasi-step functions, the {L^{p,q}} norms barely depend on {q} at all.

One easy corollary of the above theorem is that the {L^{p,q}} quasi-norm is indeed a quasi-norm, and in particular that

\displaystyle  \| f + g \|_{L^{p,q}} \lesssim_{p,q} \| f \|_{L^{p,q}} + \| g \|_{L^{p,q}} \ \ \ \ \ (15)

for any {f,g \in L^{p,q}}; this can be seen for instance by using the equivalence of (i) and (iii). Another easy consequence is that the simple functions are dense in {L^{p,q}}.

Exercise 32
  • (i) Suppose that {\| \|} is a quasinorm on some function space {X} with quasitriangle inequality constant {C \geq 1}, thus

    \displaystyle  \|f+g\| \leq C (\|f\| + \|g\|)

    for all {f,g \in X}. Let {p} be such that {(2C)^p = 2}. Show that

    \displaystyle  \| \sum_{n=1}^N f_n \| \leq 2 C_0 (\sum_{n=1}^N \|f_n\|^p)^{1/p}.

    (Hint: first show that if {\|f\|^p, \|g\|^p \leq 2^k}, then {\|f+g\|^p \leq 2^{k+1}}. Then iterate this carefully to show that if {\|f_n\|^p \leq 2^{k_n}} and {\sum_{n=1}^N 2^{k_n} \leq 2^k}, then {\| \sum_{n=1}^N f_n \|^p \leq 2^{k}}.)
  • (ii) Suppose that a sequence {f_n \in L^{p,q}} for {n \geq 0} obeys an exponential decay bound of the form

    \displaystyle  \|f_n\|_{L^{p,q}} \leq A 2^{-\varepsilon n}

    for some {A, \varepsilon > 0} and all {n \geq 0}. Show that the series

    \displaystyle  \sum_{n=0}^\infty f_n

    converges in {L^{p,q}}, with

    \displaystyle  \| \sum_{n=0}^\infty f_n \|_{L^{p,q}} \lesssim_{p,q,A,\varepsilon} 1.

Exercise 33 For each integer {n}, let {f_n} be a quasi-step function of height {2^n} and width {W_n} for some {W_n > 0}. Show that

\displaystyle  \| \sum_n |f_n| \|_p \sim_p \| 2^n W_n^{1/p} \|_{l^p_n({\bf Z})}

for all {0 < p < \infty}. If instead {f_n} is a quasi-step function of height {H_n} and width {2^n} for some {H_n > 0}, show that

\displaystyle  \| \sum_n |f_n| \|_p \sim_p \| H_n 2^{n/p} \|_{l^p_n({\bf Z})}

for all {0 < p < \infty}. What goes wrong when we remove the absolute values on the {f_n}? (This can be repaired by replacing the powers of {2} with powers of a sufficiently large constant (depending on the implied constant in the definition of a quasi-step function) – why?)

A particularly useful consequence of the above theorem is a Hölder inequality for Lorentz spaces, due to O’Neil.

Theorem 34 (Hölder’s inequality in Lorentz spaces) If {0 < p_1,p_2,p < \infty} and {0 < q_1,q_2,q \leq \infty} obey {\frac{1}{p} = \frac{1}{p_1} + \frac{1}{p_2}} and {\frac{1}{q} = \frac{1}{q_1} + \frac{1}{q_2}} then

\displaystyle  \|fg \|_{L^{p,q}} \lesssim_{p_1,p_2,q_1,q_2} \|f\|_{L^{p_1,q_1}} \|g\|_{L^{p_2,q_2}}

whenever the right-hand side norms are finite.

Proof: We may normalise {\|f\|_{L^{p_1,q_1}} = \|g\|_{L^{p_2,q_2}} = 1}, and drop the dependence of the implied constants on {p_1,p_2,q_1,q_2} for brevity. By the equivalence of (i) and (v) in Theorem 28 we may dominate {|f| \leq \sum_n H_n 1_{E_n}} and {|g| \leq \sum_n H'_n 1_{E'_n}} where {\mu(E_n), \mu(E'_n) \lesssim 2^n} and

\displaystyle  \| H_n 2^{n/p_1} \|_{l^{q_1}_n}, \| H'_n 2^{n/p_2} \|_{l^{q_2}_n} \lesssim 1.

Then we have

\displaystyle  |fg| \leq \sum_k \sum_n H_n H'_{n+k} 1_{E_n \cap E'_{n+k}}.

By the quasi-triangle inequality and monotonicity it suffices to show that

\displaystyle  \| \sum_{k \geq 0} \sum_n H_n H'_{n+k} 1_{E_n \cap E'_{n+k}} \|_{L^{p,q}} \lesssim 1

and

\displaystyle  \| \sum_{k < 0} \sum_n H_n H'_{n+k} 1_{E_n \cap E'_{n+k}} \|_{L^{p,q}} \lesssim 1.

By symmetry it suffices to consider the {k \geq 0} component. Here we observe that {E_n \cap E'_{n+k}} has measure at most {2^n}, so by the equivalence of (i) and (v) in Theorem 28

\displaystyle  \| \sum_n H_n H'_{n+k} 1_{E_n \cap E'_{n+k}} \|_{L^{p,q}} \lesssim \| H_n H'_{n+k} 2^{n/p} \|_{l^q_n}.

But by the ordinary Hölder inequality

\displaystyle  \| H_n H'_{n+k} 2^{n/p} \|_{l^q_n} \leq \| H_n 2^{n/p_1} \|_{l^{q_1}_n} \| H'_{n+k} 2^{n/p_2} \|_{l^{q_2}_n};

shifting the second {n} by {k} we conclude

\displaystyle  \| \sum_n H_n H'_{n+k} 1_{E_n \cap E'_{n+k}} \|_{L^{p,q}} \lesssim 2^{-k/p_2}.

The claim now follows from Exercise 32. \Box

One corollary of this Hölder inequality is that {L^{p,q}} functions are absolutely integrable on sets of finite measure whenever {p > 1}.

Now we consider dual formulations of the {L^{p,q}} norms. The case {q=\infty} is fairly straightforward:

Exercise 35 (Dual formulation of weak {L^p}) Let {1 < p \leq \infty}. Then for every {f} in {L^{p,\infty}(X,d\mu)}, we have

\displaystyle \|f\|_{L^{p,\infty}(X,d\mu)}

\displaystyle  \sim_p \sup \{ \mu(E)^{-1/p'} |\int_X f 1_E\ d\mu|: 0 < \mu(E) < \infty \} \ \ \ \ \ (16)

Also show that the hypothesis {f \in L^{p,\infty}(X,d\mu)} can be dropped if one instead assumes {f} to be non-negative.

The right-hand side of (16) is clearly a semi-norm at least on {f}. This leads in particular to a quasi-triangle inequality

\displaystyle  \| f_1 + \ldots + f_N \|_{L^{p,\infty}(X,d\mu)} \sim_p \|f_1\|_{L^{p,\infty}(X,d\mu)} + \ldots + \|f_N\|_{L^{p,\infty}(X,d\mu)}

for any {f_1,\ldots,f_N \in L^{p,\infty}(X,d\mu)}.

Remark 36 It is worth comparing (16) to (4). In (4), one takes the inner product of {f} against all {L^{p'}}-normalised functions, and the worst inner product becomes the {L^p} norm. In (16), one only takes the inner product of {f} against the {L^{p'}}-normalised step functions {\mu(E)^{-1/p'} 1_E}. This is fully consistent with the fact that the {L^p} norm is stronger than the weak {L^p} norm.

Exercise 35 can be rephrased as follows: if {f \in L^{p,\infty}} for some {1 < p < \infty} and {A > 0}, then the following two statements are equivalent (up to changes in the implied constants):

  • {\|f\|_{L^{p,\infty}} \lesssim_p A}.
  • {\int_X f 1_E\ d\mu = O_p( A \mu(E)^{1/p'} )} for all sets {E} of finite measure.

Unfortunately this equivalence breaks down at {p=1} or below (consider for instance the weak {L^p} function {|x|^{-d/p}} on {{\bf R}^d}, which is not even locally integrable when {p \leq 1}). However, one does have a substitute:

Exercise 37 Let {0 < p < \infty}, {0 < A < \infty}, and {f \in L^{p,\infty}(X,d\mu)}. Show that the following are equivalent (up to changes in the implied constant):
  • {\|f\|_{L^{p,\infty}(X,d\mu)} \lesssim_p A}.
  • For every set {E} of finite measure, there exists a subset {E'} of {E} with {\mu(E') \geq \frac{1}{2} \mu(E)} such that

    \displaystyle  \int_X f 1_{E'}\ d\mu = O( \mu(E)^{1/p'} )

    (in particular, we assert that the integral on the left-hand side is absolutely integrable).
Hint: It may be instructive to work out the example {f(x) = |x|^{-d/p}} on {{\bf R}^d} by hand to get a sense of what is going on; this should suggest how to prove things in general. The proof is slightly simpler in the case when {f} is non-negative, so you may want to try that case first. Comment on how this result implies Exercise
35 (or its equivalent version discussed shortly afterwards) when {p > 1}.

Exercise 38 Let {f_1,\ldots,f_N} be functions with {N \geq 2}. Show that

\displaystyle  \| f_1 + \ldots + f_N \|_{L^{1,\infty}} \lesssim \log N ( \|f_1\|_{L^{1,\infty}} + \ldots + \| f_N \|_{L^{1,\infty}} ).

thus the weak {L^{1,\infty}} quasinorm only fails to be a norm “by a logarithm”. Show with an example that the {\log N} cannot be removed.

For more general {L^{p,q}} spaces, we have

Theorem 39 (Dual characterisation of {L^{p,q}}) Let {1 < p < \infty} and {1 \leq q \leq \infty}. Then for any {f \in L^{p,q}},

\displaystyle  \|f\|_{L^{p,q}} \sim_{p,q} \sup \{ |\int_X f \overline{g}\ d\mu|: \|g\|_{L^{p',q'}} \leq 1 \}.

Again, the hypothesis {f \in L^{p,q}} can be dropped if {f} is non-negative and {g} is restricted to also be non-negative.

Thus the {L^{p,q}} quasi-norm is in fact equivalent to a norm when {1 < p < \infty} and {q \geq 1}. In particular, weak {L^p} is equivalent to a normed space when {1 < p < \infty}. (For {p=1}, weak {L^p} fails to be normable “by a logarithm”; see Q3.) As with other dual characterisations, one can restrict {g} to a dense subclass of {L^{p',q'}}, for instance simple functions with finite measure support.

Proof: To obtain the {\gtrsim_{p,q}} part of this theorem, we simply estimate

\displaystyle  |\int_X f \overline{g}\ d\mu| \leq \|fg\|_{L^1} = \|fg\|_{L^{1,1}}

and use Theorem
34. To obtain the {\lesssim_{p,q}} part, we normalise {\|f\|_{L^{p,q}} = 1}. It then suffices by homogeneity to find {g} with {\|g\|_{L^{p',q'}} \leq 1} and {\int_X f \overline{g}\ d\mu \gtrsim 1}.

The case {q=\infty} follows from (16), so let us take {q < \infty}. We will just give the proof in the case {q \geq p}; the case {p<q} is trickier, and a partial argument is given in the exercises. By the equivalence of (i) and (ii) in Theorem 28 may write {f = \sum_m f_m} where {f_m} is a quasi-step function of height {H_n} and width {W_m} with disjoint supports such that the sequence {a_m := 2^m W_m^{1/p}} has an {l^q_m} norm of {\sim_{p,q} 1}. Now take

\displaystyle  g := \sum_m g_m

where

\displaystyle  g_m := a_m^{q-p} |f_m|^{p-2} f_m

adopting the obvious convention that {|f_m|^{p-2} f_m = 0} when {f_m=0}. Then (because of the disjoint supports)

\displaystyle  |\int_X f \overline{g}| = \sum_m \int_X a_m^{q-p} |f_m|^p.

But since {f_m} has height {2^m} and width {W_m}, {\int_X |f_m|^p \sim_p 2^{mp} W_m = a_m^p} and so

\displaystyle  |\int_X f \overline{g}| \sim_p \sum_m a_m^q \sim_{p,q} 1.

To conclude it will suffice to show that

\displaystyle  \|\sum_m g_m \|_{L^{p',q'}} \lesssim_{p,q} 1.

If {E_m} is the support of {f_m}, then we have the pointwise bound

\displaystyle  g_m \lesssim_{p,q} a_m^{q-p} 2^{m(p-1)} 1_{E_m}

and the measure bound {\mu(E_m) \lesssim_{p,q} W_m = 2^{-mp} a_m^p}.

At this point we would like to apply Theorem 28, but neither the height nor width of {g_m} is necessarily a power of {2}. But we can remedy this by introducing the modified heights

\displaystyle  H_m := \sup_{k \geq 0} a_{m-k}^{q-p} 2^{m(p-1)} 2^{-k(p-1)/2}.

We have {H_{m+1} \geq 2^{(p-1)/2} H_m}, and so the {H_m} increase geometrically. It then suffices to show that

\displaystyle  \| \sum_m H_m 1_{E_m} \|_{L^{p',q'}} \lesssim_{p,q} 1.

By refining the {m} by a constant factor we can make each {H_m} at least twice as large as the previous, and so by applying the equivalences of (i) and (iii) in Theorem 28 and the triangle inequality it suffices to show that

\displaystyle  \| H_m \mu(E_m)^{1/p'} \|_{l^{q'}_m} \lesssim_{p,q} 1

which we expand using our bound on {\mu(E_m)} as

\displaystyle  \| a_m^{p-1} \sup_{k \geq 0} a_{m-k}^{q-p} 2^{-k(p-1)/2} \|_{l^{q'}_m}.

But from Hölder’s inequality (using the hypothesis {q-p \geq 0}) and the {l^q} bound on {a_m} we have

\displaystyle  \| a_m^{p-1} a_{m-k}^{q-p} 2^{-k(p-1)/2} \|_{l^{q'}_m} \lesssim_{p,q} 2^{-k(p-1)/2};

summing this using the triangle inequality (and estimating the supremum by a sum) we obtain the claim.

The case when {f,g} are restricted to be non-negative can be deduced from the above result and a monotone convergence argument (representing {f} as a monotone limit of simple functions of finite measure support) which we leave as an exercise to the reader. \Box

Exercise 40 A measure space {(X,\mu)} is said to be non-atomic if, for every measurable set {A \subset X} with {\mu(A) > 0}, there exists a measurable subset {B \subset A} such that {0 < \mu(B) < \mu(A)}.
  • (i) (Sierpinski’s theorem) Show that if {(X,\mu)} is non-atomic, {A \subset X} is measurable, and {0 \leq b \leq \mu(A)}, then there exists a measurable subset {B} of {A} such that {\mu(B) = b}. (You may find it convenient to use Zorn’s lemma.)
  • (ii) Show that if {(X,\mu)} is non-atomic, then the decomposition in Theorem 28(iv) can be chosen so that for each {n}, either {H_{n+1}} vanishes, or {f_n} has support of measure {\sim 2^n} (not just bounded by {2^n}).
  • (iii) Establish Theorem 28 in the case that {q < p} and {(X,\mu)} is non-atomic, by using the modification of Theorem 28 indicated in the previous part of the exercise.
The duality can also be established for measure spaces that contain atoms, but this requires a more careful analysis that treats large atoms separately.

— 3. Orlicz spaces (Optional) —

So far we have studied the Lebesgue spaces {L^p}, together with the more general Lorentz spaces {L^{p,q}}, which includes weak {L^p} as a special case. These spaces are all rearrangement-invariant and monotone. There is a different generalisation of the Lebesgue spaces {L^p}, the Orlicz spaces {\Phi(L)}, which are also rearrangement-invariant and monotone, and which are occasionally useful. (There is a common generalisation of both, the Lorentz-Orlicz spaces, but these occur very rarely in applications.)

The motivation for Orlicz spaces starts with the trivial observation that if {1 \leq p < \infty}, then

\displaystyle  \|f\|_{L^p} \leq 1 \hbox{ if and only if } \int_X |f|^p\ d\mu \leq 1.

Inspired by this, we generalise by letting {\Phi: {\bf R}^+ \rightarrow {\bf R}^+} be a function (with some additional properties to be selected shortly) and ask if we can find a norm {\| f \|_{\Phi(L)}} which obeys the property

\displaystyle  \|f\|_{\Phi(L)} \leq 1 \hbox{ if and only if } \int_X \Phi(|f|)\ d\mu \leq 1. \ \ \ \ \ (17)

Since norms need to be homogeneous, this would imply

\displaystyle  \|f\|_{\Phi(L)} \leq A \hbox{ if and only if } \int_X \Phi(|f|/A)\ d\mu \leq 1

for all {A > 0}. In particular, if {A < A'}, then we need

\displaystyle  \int_X \Phi(|f|/A)\ d\mu \leq 1 \hbox{ implies } \int_X \Phi(|f|/A')\ d\mu \leq 1.

To ensure this property it is thus very natural to require that {\Phi} be increasing. Also to deal with the zero norm case one typically requires {\Phi(0)=0}.

Next, in order for {\Phi(L)} to be a norm, the unit ball {\{ f: \| f \|_{\Phi(L)} \leq 1 \}} needs to be convex. Looking at (17), we see that this will indeed be the case when {\Phi} is itself convex. (Note that the proof of (5) was a special case of this argument).

We can put all the above discussion together and conclude: if {\Phi: {\bf R}^+ \rightarrow {\bf R}^+} is increasing and convex with {\Phi(0)=0}, then the norm

\displaystyle  \| f \|_{\Phi(L)} := \inf \{ A > 0: \int_X \Phi(|f|/A)\ d\mu \leq 1 \}

is a norm on the space {\Phi(L) := \{ f: \|f\|_{\Phi(L)} < \infty \}}.

As discussed above, the {L^p} spaces for {1 \leq p < \infty} are examples of Orlicz spaces with {\Phi(x) := x^p}. The space {L^\infty} is not really an Orlicz space, but can be viewed the limiting case where {\Phi(x)} is infinite for {x > 1} and zero for {x \leq 1} (or more informally, {\Phi(x) = x^{+\infty}}). Aside from the Lebesgue spaces, the most common Orlicz spaces which appear are

  • The space {L \log L}, defined as the Orlicz space with {\Phi(x) := x \log (2+x)};
  • The space {e^L}, defined as the Orlicz space with {\Phi(x) := e^x - 1};
  • The space {e^{L^2}}, defined as the Orlicz space with {\Phi(x) := e^{x^2} - 1}.

The correction factors of {2} and {1} in the above functions should not be taken too seriously; note that if two functions {\Phi, \tilde \Phi} are comparable then their Orlicz norms are comparable also; a little more generally, if {\Phi \lesssim \tilde \Phi}, then {\|f\|_{\Phi(L)} \lesssim \| f \|_{\tilde \Phi(L)}}. It is the behaviour of {\Phi(x)} for large values of {x} which is the most important, although when {X} has infinite measure the behaviour at small values of {x} is also relevant.

Exercise 41 If {X} has finite measure, verify the relation

\displaystyle  \| f \|_{\Phi(L)} \sim_{\Phi, \mu(X)} \| f \|_{1+\Phi(L)}

which is another indication of the irrelevance of the low values of {\Phi} in the finite measure case.

The final fact about Orlicz spaces that we give here is the duality relation. Suppose that {\Phi: {\bf R}^+ \rightarrow {\bf R}^+} is increasing, convex, and is also superlinear in the sense that {\lim_{x \rightarrow +\infty} \Phi(x)/x = +\infty}. We can then define the Young dual {\Psi: {\bf R}^+ \rightarrow {\bf R}^+} of {\Phi} by the formula

\displaystyle  \Psi(y) := \sup \{ xy - \Phi(x): x \in {\bf R}^+ \};

the hypothesis that {\Phi} is superlinear ensures that this function is well-defined. We may equivalently define {\Psi(y)} to be the smallest function for which one has the inequality

\displaystyle  xy \leq \Phi(x) + \Psi(y) \hbox{ for all } x, y \in {\bf R}^+. \ \ \ \ \ (18)

Exercise 42 If {1 < p < \infty}, show that the Young dual of {\Phi(x) = x^p} is {\Psi(y) = \frac{p^{p'/p}}{p'} y^{p'}}. Show also that the Young dual of {\Phi(x) = x \log(2+x)} takes the form {\Psi(y) \sim e^y} for {y > 1}. What does the Young duals of {e^x - 1} and {e^{x^2}-1} look like?

One can easily verify that {\Psi} is also increasing, convex, and super-linear, and so the Orlicz norm {\| \|_{\Psi(L)}} makes sense. From (18) and the triangle inequality it is immediate that

\displaystyle  |\int_X f \overline{g}\ d\mu| \leq 2 \hbox{ whenever } \|f\|_{\Phi(L)}, \|g\|_{\Psi(L)} \leq 1

and hence by homogeneity we obtain the duality relation

\displaystyle  |\int_X f \overline{g}\ d\mu| \leq 2 \|f\|_{\Phi(L)} \|g\|_{\Psi(L)}

whenever {f \in \Phi(L)} and {g \in \Psi(L)}.

Exercise 43 Establish the more precise relationship

\displaystyle  \| f \|_{\Phi(L)} \sim \sup \{ |\int_X f \overline{g}\ d\mu| : \|g\|_{\Psi(L)} \leq 1 \}.

Exercise 44 Let {\Phi} be superlinear, increasing, convex, and vanishing at the origin.
  • (i) Show that the Young dual {\Psi} of {\Phi} is also superlinear, increasing, convex, and vanishing at the origin.
  • (ii) Show that {\Phi} is the Young dual of {\Psi}.
(It may help to view things geometrically, and in particular understanding {\Psi} as parameterising the support lines of the graph of the convex function {\Phi}.)

Exercise 45 When {X} has finite measure, show that the spaces {L \log L} and {e^L} are dual to each other. What is the dual to {e^{L^2}}?

Exercise 46 Let {f} be a function on a measure space {X} of bounded measure {\mu(X) = O(1)}. Show that the following are equivalent (up to changes in the implied constants):
  • (i) {\|f\|_{e^L} = O(1)}.
  • (ii) {\|f\|_{L^{p,\infty}} = O(p)} for all {1 \leq p < \infty}.
  • (iii) {\|f\|_{L^p} = O(p)} for all {1 \leq p < \infty}.
Hint: You may find the Taylor expansion for {e^x}, together with the obvious bounds {(k/2)^{k/2 - 1} \leq k! \leq k^k} for integer {k} (or Stirling’s formula, if you know what that is), to be useful.

Exercise 47 Obtain the analogue of Exercise 46 for the Orlicz space {e^{L^2}}.

— 4. Real interpolation —

So far we have only considered functions {f} on a single measure space {X = (X, {\mathcal B}_X, \mu_X)}. Now we shall consider operators {T} which take functions on one measure space {X = (X, {\mathcal B}_X, \mu_X)} to functions on another measure space {Y = (Y, {\mathcal B}_Y, \mu_Y)}; the study of such operators is in fact a major focus of harmonic analysis. Ultimately we want to extend {T} to a standard normed vector space such as {L^p(X)}, but in practice one has to initially first restrict attention to a dense subspace of functions, such as simple functions or test functions.

We are primarily interested in linear operators, thus {T(cf) = c Tf} and {T(f+g) = Tf + Tg}. But it is also worth considering the more general sublinear operators, in which

\displaystyle  |T(cf)| = |c| |Tf|

and we have the pointwise estimate

\displaystyle  |T(f+g)| \leq |Tf| + |Tg|.

Apart from the linear operators, the next most important example of a sublinear operator is a maximal operator

\displaystyle  Tf(x) := \sup_n |T_n f(x)|

where {T_n} are a collection (possibly countably or uncountably infinite, though in the latter case one has to take some care in ensuring measurability) of linear or sublinear operators. The third most important example is a square function such as

\displaystyle  Tf(x) := (\sum_n |T_n f(x)|^2)^{1/2}.

More generally, one can consider a family {T_y} of operators indexed by some parameter {y}, and take {Tf(x)} to be the norm in the {y} variable of {T_y f} in some suitable norm. But the above three examples of linear operators, maximal operators, and square functions already cover the vast majority of applications.

Let {0 < p,q \leq \infty} be exponents, and let {T} be sublinear. Let us define the following concepts.

  • We say that {T} is strong-type {(p,q)} (or simply type {(p,q)}) if we have a bound

    \displaystyle  \| Tf \|_{L^q(Y)} \lesssim_{T,p,q} \| f \|_{L^p(X)}

    for all {f} in {L^p}, or in a dense sub-class thereof. Note that in the latter case there is a unique extension to all of {L^p}.

  • If {q < \infty}, we say that {T} is weak-type {(p,q)} if we have a bound

    \displaystyle  \| Tf \|_{L^{q,\infty}(Y)} \lesssim_{T,p,q} \| f \|_{L^p(X)}

  • We say that {T} is restricted strong-type {(p,q)} if we have a bound

    \displaystyle  \| T f \|_{L^{q}(Y)} \lesssim_{T,p,q} H W^{1/p} \ \ \ \ \ (19)

    for all sub-step functions of height {H} and width {W}. In particular, we have

    \displaystyle  \| T 1_E \|_{L^{q}(Y)} \lesssim_{T,p,q} \mu(E)^{1/p}. \ \ \ \ \ (20)

    (Conversely, we can deduce (19) from (20) using tricks such as those in Remark 27.)

  • If {q < \infty}, we say that {T} is restricted weak-type {(p,q)} if we have a bound

    \displaystyle  \| T f \|_{L^{q,\infty}(Y)} \lesssim_{T,p,q} H W^{1/p} \ \ \ \ \ (21)

    for all sub-step functions {f} of height {H} and width {W}. In particular, we have

    \displaystyle  \| T 1_E \|_{L^{q,\infty}(Y)} \lesssim_{T,p,q} \mu(E)^{1/p}. \ \ \ \ \ (22)

Clearly, whenever {p,q} are fixed, strong-type implies weak-type and restricted strong-type, either of which imply restricted weak-type. In most applications, it is the strong-type bounds which are desired; however, we shall see in this section that the real interpolation method allows us to deduce strong-type bounds from weak-type, or even restricted weak-type bounds, as long as the strong-type bounds are an interpolant between the restricted weak-type bounds. This can be a very useful strategy, because (as we shall see in next week’s notes) weak-type or restricted weak-type bounds are easier to prove than strong-type estimate.

Let us first make a mild (and qualitative) assumption, namely that the form

\displaystyle  \langle |Tf|, |g| \rangle := \int_Y |Tf| |g|\ d\nu \ \ \ \ \ (23)

is well-defined whenever {f, g} are simple functions with finite measure support. This is for instance the case if {T} is of restricted type {(p,q)} for some {0 < p < \infty} and {1 \leq q \leq \infty}, or restricted weak-type {(p,q)} for some {0 < p < \infty} and {1 < q < \infty}; thus in practice this assumption is easily satisfied. We observe that this form is non-negative, homogeneous and sublinear in both {f} and {g}:

\displaystyle  \begin{array}{rl}  \langle |Tcf|, |g| \rangle = \langle |Tf|, |cg| \rangle &= |c| \langle |Tf|, |g| \rangle \\ \langle |T(f_1+f_2)|, |g| \rangle &\leq \langle |T f_1|, |g| \rangle + \langle |T f_2|, |g| \rangle \\ \langle |Tf|, |g_1+g_2| \rangle &\leq \langle |T f|, |g_1| \rangle + \langle |T f|, |g_2| \rangle. \end{array}

The form {\langle |Tf|, |g| \rangle} turns out to be a convenient way to understand the various types of {T}, and the ability to decompose both {f} and {g} independently is very useful in establishing interpolation type results. There is a near-symmetry between {f} and {g} here, if we could somehow take an “adjoint” of the operator {T}, but we will not explicitly exploit this symmetry here as it is not always available for sublinear operators (though the “duality” or “adjoint” trick is undoubtedly very powerful in the important linear case).

Let us look in particular at the form (23) applied to indicator functions, thus we look at the quantity {\langle |T1_E|, 1_F \rangle} for {E \subset X} and {F \subset Y} of finite measure. Now suppose that {T} had some strong type {(p,q)} bound for some {0 < p < \infty} and {1 \leq q \leq \infty}, say

\displaystyle  \| Tf \|_{L^q(Y)} \lesssim_{p,q} A \|f\|_{L^p(X)}

for some {A > 0}. Then in particular

\displaystyle  \| T1_E \|_{L^q(Y)} \lesssim_{p,q} A \mu(E)^{1/p}

and hence by Hölder’s inequality

\displaystyle  \langle |T1_E|, 1_F \rangle \lesssim_{p,q} \mu(E)^{1/p} \nu(F)^{1/q'}. \ \ \ \ \ (24)

Actually it is clear that strong type {(p,q)} is too much of an assumption; restricted type {(p,q)} would have sufficed for the conclusion. If {q > 1}, we can relax things further to restricted weak-type:

Exercise 48 Let {0 < p \leq \infty}, {1 < q \leq \infty}, and {A > 0}. Let {T} be a sublinear operator such that the form (23) is well-defined. Then the following are equivalent up to changes in the implied constant:
  • {T} is restricted weak-type {(p,q)} with constant {A}, in the sense that

    \displaystyle  \| T f \|_{L^{q,\infty}(Y)} \lesssim_{p,q} A H W^{1/p} \ \ \ \ \ (25)

    for all simple sub-step functions {f} of height {H} and width {W}.
  • For all {E \subset X}, {F \subset Y} of finite measure, we have the bound

    \displaystyle  \langle |T1_E|, 1_F \rangle \lesssim_{p,q} A \mu(E)^{1/p} \nu(F)^{1/q'}

(Hint: use (16) and Remark 27.)

This already gives a simple version of real interpolation:

Corollary 49 (Baby real interpolation) Let {T} be a sublinear operator such that the form (23) is well-defined. Let {0 < p_0,p_1 \leq \infty}, {1 < q_0,q_1 \leq \infty} and {A_0,A_1 > 0} be such that {T} is restricted weak-type {(p_i,q_i)} with constant {A_i} (in the sense of (25)) for {i=0,1}. Then {T} is also restricted weak-type {(p_\theta,q_\theta)} with constant {A_\theta} for {0 \leq \theta \leq 1}, where

\displaystyle  \frac{1}{p_\theta} := \frac{1-\theta}{p_0} + \frac{\theta}{p_1}; \quad \frac{1}{q_\theta} := \frac{1-\theta}{q_0} + \frac{\theta}{q_1}; A_\theta := A_0^{1-\theta} A_1^\theta \ \ \ \ \ (26)

and the implied constant depends on {p_0,p_1,q_0,q_1}.

Indeed, all we are using here is the obvious algebraic observation that if {X \lesssim Y_0} and {X \lesssim Y_1}, then {X \lesssim Y_\theta := Y_0^{1-\theta} Y_1^\theta} for all {0 \leq \theta \leq 1}.

The above corollary has two defects. Firstly, it can only conclude restricted weak-type rather than strong type. Secondly, the restriction {q_0, q_1 > 1} is inconvenient for many applications, since weak {L^{1,\infty}} bounds are actually rather common. To address the second concern, we have the following variant of Exercise 48:

Exercise 50 Let {0 < p \leq \infty}, {0 < q \leq \infty}, and {A > 0}. Let {T} be a sublinear operator such that the form (23) is well-defined. Then the following are equivalent up to changes in the implied constant:
  • {T} is restricted weak-type {(p,q)} with constant {A}, in the sense that (25) holds.
  • For all {E \subset X}, {F \subset Y} of finite non-zero measure, there exists {F' \subset F} with {\mu(F') \geq \frac{1}{2} \mu(F)} such that

    \displaystyle  \langle |T1_E|, 1_{F'} \rangle \lesssim_{p,q} A \mu(E)^{1/p} \nu(F)^{1/q'}

Hint: use Exercise 37.

Corollary 51 The hypothesis {q_0,q_1 > 1} in Corollary 49 can be weakened to {q_0,q_1 > 0}.

Now we can give a significantly more useful real interpolation theorem, which can interpolate between restricted weak-type estimates to obtain strong-type estimates.

Theorem 52 (Marcinkiewicz interpolation theorem) Let {T} be a sublinear operator such that the form (23) is well-defined. Let {0 < p_0,p_1 \leq \infty}, {0 < q_0,q_1 \leq \infty} and {A_0,A_1 > 0} be such that {T} is restricted weak-type {(p_i,q_i)} with constant {A_i} (in the sense of (25)) for {i=0,1}. Suppose also that {p_0 \neq p_1} and {q_0 \neq q_1}. Then for any {0 < \theta < 1} and {1 \leq r \leq \infty} with {q_\theta > 1} we have

\displaystyle  \| Tf \|_{L^{q_\theta,r}(Y)} \lesssim_{p_0,p_1,q_0,q_1,r,\theta} A_\theta \| f \|_{L^{p_\theta,r}(X)}

for all simple functions {f} with finite measure support, where {p_\theta, q_\theta, A} were defined in (26). In particular, if {q_\theta \geq p_\theta}, then {T} is strong-type {(p_\theta,q_\theta)} with constant {O_{p_0,p_1,q_1,\theta}(A_\theta)}.

Proof: To simplify the notation let us suppress the dependence on {p_0,p_1,q_0,q_1,r,\theta}. Observe that the statement and conclusion of the theorem have several homogeneity symmetries. The most obvious one is that we can multiply {T} (and the {A_\theta}) by an arbitrary constant, but we may also multiply the measure {\mu} by a constant {C} (and the {A_\theta} by the constant {C^{-1/p_\theta}}); similarly we may multiply {\nu} by {C} and {A_\theta} by {C^{1/q_\theta}}). Using these symmetries we may normalise {A_0=A_1=1} (and hence {A_\theta = 1} for all {\theta}). Our task is then to show that

\displaystyle  \| Tf \|_{L^{q_\theta,r}(Y)} \lesssim \| f \|_{L^{p_\theta,r}(X)}

for all simple functions {f} with finite measure support.

Using Theorem 39 (and the hypothesis {q_\theta > 1}), this is equivalent to showing that

\displaystyle  \langle |Tf|, |g| \rangle \lesssim \| f \|_{L^{p_\theta,r}(X)} \| g \|_{L^{q'_\theta,r'}(Y)} \ \ \ \ \ (27)

for all simple functions {f,g} of finite measure support.

We currently have {q_0,q_1 > 0} and {q_\theta > 1}. By using Corollary 51 to bring the restricted weak-type exponents {(p_0,q_0)} and {(p_1,q_1)} a bit closer to {(p_\theta,q_\theta)} we may assume that {q_0,q_1 > 1} as well. Applying Exercise 48 we conclude that

\displaystyle  \langle |T1_E|, 1_F \rangle \lesssim |E|^{1/p_i} |F|^{1/q'_i}

for all sets {E, F} of finite measure and {i=0,1}. From Remark 27 and sublinearity we conclude that

\displaystyle  \langle |Tf|, |g| \rangle \lesssim H W^{1/p_i} H' (W')^{1/q'_i}

whenever {f,g} are sub-step functions of height and widths {H,H'} and {W,W'} respectively. We can of course pick the better of the two estimates, leading to

\displaystyle  \langle |Tf|, |g| \rangle \lesssim H H' \min_{i=0,1}( W^{1/p_i} (W')^{1/q'_i} ). \ \ \ \ \ (28)

Now we can return to proving (27). By homogeneity we may normalise

\displaystyle  \| f \|_{L^{p_\theta,r}(X)} = \| g \|_{L^{q'_\theta,r'}(Y)} = 1.

We then apply Theorem 28 to decompose {f = \sum_{n \in {\bf Z}} f_n}, {g = \sum_{n \in {\bf Z}} g_n} with {f_n, g_n} sub-step functions of width {2^n} and heights {H_n, H'_n} respectively, with the height bounds

\displaystyle  \| a \|_{l^r({\bf Z})}, \| b \|_{l^{r'}({\bf Z})} \lesssim 1 \ \ \ \ \ (29)

where {a,b} are the sequences

\displaystyle  a_n := H_n 2^{n/p_\theta}; \quad b_n := H'_n 2^{n/q'_\theta}.

Since {f,g} are simple functions, only finitely many of the {f_n} and {g_n} are non-zero by Remark 30. We now use sublinearity to estimate

\displaystyle  \langle |Tf|, |g| \rangle \leq \sum_{n,m} \langle |Tf_n|, |g_m| \rangle

and then use (28) to obtain

\displaystyle  \langle |Tf|, |g| \rangle \lesssim \sum_{n,m} H_n H'_m \min_{i=0,1} ( 2^{n/p_i} 2^{m/q'_i} ).

We can write this in terms of {a}, {b}, and reduce to showing that

\displaystyle  \sum_{n,m} a_n b_m \min_{i=0,1}( 2^{n (\frac{1}{p_i}-\frac{1}{p_\theta})} 2^{m (\frac{1}{q_\theta} - \frac{1}{q_i})} ) \lesssim 1.

Because {p_0 \neq p_1} and {q_0 \neq q_1}, and because of the definitions of {p_\theta, q_\theta}, we can write the left-hand side as

\displaystyle  \sum_{n,m} a_n b_m \min( 2^{\varepsilon(n+\alpha m)}, 2^{-\varepsilon'(n+\alpha m)} )

for some non-zero {\alpha \in {\bf R}} and some {\varepsilon,\varepsilon' > 0} depending only on {p_0,p_1,q_0,q_1,\theta}. We substitute {k := \lfloor n + \alpha m \rfloor} (thus {n = k - \lfloor \alpha m \rfloor}) and estimate this by

\displaystyle  \sum_k \min( 2^{\varepsilon k}, 2^{-\varepsilon k} ) \sum_{m} a_{k - \lfloor \alpha m} b_m.

By (29) and Hölder we see that the inner sum is {O(1)} uniformly in {k}, and the claim follows.

Finally, if we specialise {r = q_\theta} and recall that the {L^{p_\theta,q_\theta}} norm will be dominated by the {L^{p_\theta}} norm for {p_\theta \geq q_\theta}, the last claim of the theorem follows. \Box

There are many other variations on the real interpolation method, for instance an extension to multilinear operators, or to other function spaces. However, the basic method of proof is still the same: dualise, decompose all inputs, estimate each term as optimally as one can, and then sum.

One can illustrate the real interpolation method graphically using type diagrams. One plots all points {(\frac{1}{p}, \frac{1}{q})} where the operator {T} is strong-type {(p,q)} or restricted weak-type {(p,q)}. Ignoring some of the technical hypotheses, the above interpolation theorems then essentially assert that the restricted weak-type diagram and strong-type diagrams are both convex, with the latter contained in the former. Furthermore, if two points lie in the former, then the open interval connecting them lies in the latter. Determining the type diagrams of various operators {T} is a basic task of harmonic analysis, as it conveys a lot of information as to how {T} transforms the width and height of functions.

There is a different interpolation method, the complex interpolation method, which offers similar results to the real interpolation method but with some slight differences. On the plus side, the complex interpolation method gives sharper bounds, and more importantly can handle the case where the operator {T} itself varies (analytically) with the interpolation parameter. On the minus side, the method cannot upgrade weak or restricted weak-type estimates to strong-type estimates.

In the next set of notes we shall present several applications of the real interpolation method.

— 5. Miscellaneous exercises —

Exercise 53 (Loomis-Whitney inequality) Let {d \geq 2}, let {X_1,\ldots,X_d} be measure spaces, and for {i=1,\ldots,d} let {f_i \in L^p(\prod_{1 \leq j \leq d: j \neq i} X_j)} for some {0 < p \leq \infty}. Show that the function

\displaystyle  F(x_1,\ldots,x_d) := \prod_{i=1}^d f_i( x_1,\ldots,x_{i-1},x_{i+1},\ldots,x_d)

lies in {L^{p/(d-1)}(\prod_{1 \leq j \leq d} X_j)} with the Loomis-Whitney inequality

\displaystyle  \|F\|_{L^{p/(d-1)}(\prod_{1 \leq j \leq d} X_j)} \leq \prod_{i=1}^d \| f_i \|_{L^p(\prod_{1 \leq j \leq d: j \neq i} X_j)}.

Conclude in particular the box inequality

\displaystyle  \mu_{\prod_{1 \leq j \leq d} X_j}(E) \leq (\prod_{i=1}^d \mu_{\prod_{1 \leq j \leq d: j \neq i}}( \pi_i(E) ))^{1/(d-1)}

where {E} is any subset of {\prod_{1 \leq j \leq d} X_j} and {\pi_i} is the canonical projection from {\prod_{1 \leq j \leq d} X_j} to {\prod_{1 \leq j \leq d: j \neq i} X_j}. From this, deduce the weak isoperimetric inequality

\displaystyle  |E| \lesssim_d |\partial E|^{d/(d-1)}

for any {E \subset {\bf R}^d}, where {|E|} is the Lebesgue measure of {E} and {|\partial E|} is the {d-1}-dimensional Hausdorff measure of the boundary of {E}.

Exercise 54 (Borel-Cantelli lemma for functions) Let {f_1, f_2, \ldots \in L^1(X)} be such that {\sum_{n=1}^\infty \|f_n\|_{L^1(X)} < \infty}. Show that {f_n} converges to zero pointwise almost everywhere.

Exercise 55 (Borel-Cantelli lemma for sets) Let {E_1, E_2, \ldots \subset X} be such that {\sum_{n=1}^\infty \mu(E_n) < \infty}. Show that almost every {x \in X} is contained in only finitely many of the {E_n}.

Doug Natelson — Brief science items - September 2026 edition

Several science items of interest from recent weeks.  As always, apologies for missing some, which I'm sure to have done.

  • A month ago, Premi Chandra, Piers Coleman, and Clare Yu published a really nice memorial biography of Phil Anderson, one of the great scientists and characters of 20th century physics.
  • A few days ago Dam Son, an outstanding condensed matter theorist at U Chicago, posted a paper on his site that resolves a math issue that had been lingering related to a well-known paper from back in the heyday of anyon superconductivity.  The authors had always been afraid that their approach accidentally violated a sum rule.  It turns out, the authors had just misplaced a factor of 1/2, and their approach was actually exactly right.  The really novel bit here is that the paper is written as if it is single-authored by Claude, and acknowledges Dr. Son for prompts that led to the result.  See here for an interesting twitter thread on this via Sankar Das Sarma.  
  • Tangentially related, Anthropic and Matt von Hippel achieved a new result, a 9-loop perturbation theory calculation based on N=4 super Yang-Mills theory.  These kinds of mathematically virtuosic calculations are exactly the sort of task that AI tools to which are extremely well suited now.  
  • Three weeks ago, a neat paper appeared in Science Advances.  These folks used quantum interference of a Bose-Einstein condensate to test the equivalence principle in a quantum system.  The basic idea is, you take an ultracold atomic gas in a well-defined quantum state.  You split it and toss one component of it upward, while you hold the other component of it fixed in the lab frame. (This work is a descendent of this approach, which was happening in the lab above mine in back when I was in grad school.) The upward-thrown component accumulates quantum phase as it rises in the lab gravitational field, slows to a stop, and comes back down.  There is a specific amount of phase difference between the two components predicted by the equivalence principle (which says that inertial mass and gravitational mass should be exactly the same), and that's what the authors find.
  • A couple of weeks ago, this paper appeared (shoutout to the first author, who is a Rice undergrad alum).   The authors fabricated a nanomechanical resonator made out of LiNbO3 and coupled it to a qubit to act as a measurement device.  Remarkably, through the qubit they are able to see transitions between individual vibrational quantum numbers in this many-atom mechanical resonator.  This is quite impressive, and it opens up many possibilities for transduction between different quantum degrees of freedom.  
More soon.

September 27, 2026

John Preskill — Marcus theory and Marcus practice

Rudy Marcus was 91 years old when I took a course from him. His age was, arguably, not his most remarkable quality.

The course took place during the second spring of my PhD. I was completing not only the academic year, but also my course requirements. Why not, I figured, indulge in traditional statistical mechanics? Statistical mechanics neighbors, and partially overlaps with, thermodynamics. Modern approaches to statistical mechanics include fancy toolkits imparted in physics courses: the renormalization-group method; Ginzburg–Landau theory; and the latter theory’s successor, topological order. I craved a more traditional approach, as won’t surprise readers of this blog. So I ditched the physics course list and signed up for Chemistry 166.

Rudy didn’t disappoint. He began the course with a tradition passed down from a founder of thermodynamics: Boltzmann’s H theorem. The theorem amounts to the first argument ever published for the second law of thermodynamics: every (sufficiently large) closed, isolated system’s entropy increases or remains constant; it doesn’t decrease. Rederiving Boltzmann’s H theorem, I felt as I would while holding a lace collar worn by Queen Elizabeth I: as though I were touching history. 

Rudy proceeded through more of the greatest hits in statistical mechanics. For instance, the Fokker–Planck and Langevin equations describe random processes. Examples include the subject of one of Einstein’s most famous papers: the motion of a pollen grain suspended in a liquid and buffeted by the liquid’s molecules. Chemistry 166 provided much of the bread and butter, meat and potatoes, and even tofu and edamame of a thermodynamicist’s education.

Rudy advanced step by step, rederiving every result in class. He recommended a textbook; but I refer to my lecture notes, rather than to the text, to refresh my memory about course topics nowadays. When I encounter complex integrals (a certain type of calculus) in statistical mechanics, I think back to Rudy’s treatment of them. The course also crystallized my understanding of the fluctuation–dissipation theorem, which describes systems perturbed slightly out of equilibrium; examples include a magnet subjected to a weak magnetic field.1 For the course’s final project, I studied quantum master equations—which featured in a Quantum Frontiers post afterward.

No less than his explanations of science, Rudy’s enthusiasm for science impressed itself upon me. Again, he taught Chemistry 166 at age 91. He continued to collaborate on research. Whenever I saw him, even after the course ended, he asked for news about quantum information theory—which he didn’t work on. 

Two and a half years after the course, Rudy gave a speech at a dinner in honor of a couple’s engagement. The groom was pursuing a PhD in applied physics, and the bride worked as an engineer. The groom asked if Rudy had advice for young scientists. Rudy responded immediately: plenty of science remains to be done, so keep exploring.

During his speech, Rudy described his upbringing in Montreal. His high school had produced a “who’s who” of Canadians, in his words—and not because the school demanded high tuitions or offered tutors and social connections. Discrimination pushed Jewish working-class children out of other schools and into Baron Byng, where some students worked hard enough to thrive. Alumni include artists; judges; physicists; William Shatner, who played Captain Kirk in the original Star Trek series…and Rudy Marcus.

Which leads me to the quality arguably more remarkable than Rudy’s teaching Chemistry 166 at age 91: he’d won a Nobel Prize in chemistry. Unlike most chemists, Rudy worked as a theorist, rather than an experimentalist. He developed a framework now called Marcus theory. It predicts the rate at which electrons hop between molecules. I can’t tell you much more about Marcus theory because I don’t know much more about it: Rudy had enough humility that, to my memory, he mentioned his theory only once in class, in passing.

Although I know little about Marcus theory, I’ve seen much of Marcus practice. In 2022, I emailed Rudy to ask where a holiday card could reach him. The pandemic had isolated enough people that I’d determined to reach out to more acquaintances. Rudy replied, “I have been working at home, meeting twice a week on zoom with my now small research group.” He was 98 years old. His curiosity, learning, mentoring, and exploring continued.

Rudy passed away this July, five days before his 103rd birthday. I hope to learn Marcus theory someday. But I’d be even more grateful to undertake Marcus practice for even a decent fraction of the time for which he did.

1How rapidly does the system respond to the perturbation? This nonequilibrium response depends on a correlation function evaluated on an equilibrium state, according to the fluctuation–dissipation theorem. To understand why, imagine preparing the system in a thermal state, perturbing the system at an early time, and measuring the system later. What information can you extract about the perturbation? A two-time correlation function encodes this information. Imagine Fourier-transforming a two-time correlation function. The Fourier transform is proportional to a second derivative of a free energy, according to the fluctuation–dissipation theorem. (The second derivative should sound plausible because it depends on a two-time correlation function.) Second derivatives of free energies equal response functions, such as the magnetic susceptibility. Therefore, a nonequilibrium response depends on an equilibrium correlator.

Tommaso Dorigo — Flying With Gamma-Ray Counters: Another Test Or Two Of The Radiacode Zero

Flying With Gamma-Ray Counters: Another Test Or Two Of The Radiacode Zero

Yesterday I traveled all day, moving my family back to Lulea from Padova. The trip is long because there is no direct connection from Venice to Stockholm (Norwegian has one, but I prefer to avoid them as they canceled my flight recently).

Tommaso Dorigo
Categories

John Preskill — Nicole’s guide to writing and editing

Freshman year of college, I took a writing seminar from German-literature professor Ellis Shookman. Professor Shookman loved Mozart’s music, he told us early in the term. He listened to Mozart on the radio while driving from campus to Boston. Static might mar the transmission, but he could often turn up the volume and continue enjoying the program. Sometimes, the static worsened during the drive. It could worsen and worsen, until Professor Shookman’s frustration outweighed his delight at listening. He’d switch off the radio.

As Professor Shookman loved listening to Mozart’s music, he loved reading about students’ ideas. Yet static can mar a piece of writing: infelicities in grammar, structure, composition, word choice, and more. If enough infelicities obscure the writing, the frustration of reading outweighs the benefits. Professor Shookman will quit reading.

Professor Shookman marked up our essays with a blue pencil that achieved the status of legend among his students. If you’ve written a paper I’ve coauthored, you’ve probably received PDF drafts replete with green highlighting.1 A sticky note explains the reason for each highlighting: “Singular–plural mismatch.” “Active voice >> passive voice.” “Let’s clue the reader in as to this formula’s meaning before lobbing the math at them.” 

Over the past year, I’ve catalogued the suggestions I write most often on paper drafts. The comments embody principles gleaned from Strunk and White’s The Elements of Style; the Physical Review style guide; other writing guides I esteem; literature whose writing I esteem;2 collaborations with professional editors; and writing instructors, including Professor Shookman. Each section below begins with more-important principles, shading into more-nuanced ones.

Please use and disseminate these principles. Train your favorite large-language model (LLM) on them, and have the LLM critique your manuscripts. Instruct it to use green highlighting if you wish. Even if the LLM suggests fixes initially, tell it to stop offering suggestions later, so that you can devise the solutions: train not only the LLM, but also yourself. I hope to enjoy your papers as much as Professor Shookman enjoyed his sonatas.

  1. Organization
    1. Motivate your work; then, present it; and then, explain its physical significance.
    2. Begin each paragraph with a topic sentence.
    3. Begin each section, apart from the introduction and conclusion, with (i) a statement of the takeaway and (ii) an outline of the section. When outlining a section, hyperlink to each subsection. Similar guidelines concern subsections and subsubsections.
    4. Before presenting a piece of math, sketch its meaning and origin. This strategy enables the reader to understand the math as soon as they encounter it. If you throw math at the reader without introducing it, the reader will have to squint at the symbols for a while to figure out what the expression means and where it comes from.
      • Example: To calculate the average work, we substitute the Hamiltonian formula (10) into the definition (12): [equation].
    5. Most citations belong at the ends of (i) sentences and (ii) phrases concluded with commas. Put a citation elsewhere only if you have a compelling reason for doing so.
    6. Bridge each component of your writing to the next component; the next shouldn’t sound like a non sequitur.
      • Suppose that the next sentence refers to (i) a topic mentioned in the previous sentence and (ii) a new topic. Mention (i) before (ii).
        • Example: Smith et al. applied control theory to the extent possible. The attempt led to intractable equations, unlike our approach.
        • Example of a broken bridge: Smith et al. applied control theory to the extent possible. Our approach does not involve intractable equations, unlike theirs.
    7. Whenever you tell a story, tell it from start to finish, step by step. Derivations, proofs, and descriptions of experiments qualify as stories.
      • This guideline extends to descriptions of experimental setups and of mathematical objects. For example, imagine referring to an element of a subgroup of the group generated by some operators. Did you have to read the preceding sentence multiple times to process it? The sentence begins at the end of a story, then rewinds to the story’s beginning. This structure impedes understanding. The subgroup forms the context for the subgroup element, which one can’t grasp until hearing about the subgroup. The subgroup participates in a similar relationship with the group, as does the group with its generators. Therefore, one should introduce the generators, then the group, then the subgroup, and then the subgroup element.
  2. Word choice
    1. Use strong, specific words, rather than weak words.
      1. Verbs and nouns are stronger than adjectives and adverbs.
      2. Choose specific verbs (e.g., “prepare,” “evolve,” and “measure”), rather than vague, general verbs (e.g., variants of “to be” and “take,” as in “take a measurement”).
    2. Avoid statements such as “we investigate,” “we study,” and “we analyze.” Such statements don’t relate that you’ve accomplished anything. State what you’ve accomplished. Verbs such as “prove,” “test,” “confirm,” “discover,” and “find” achieve this goal.
    3. Adverbs such as “importantly” and “remarkably” pollute scientific writing with the authors’ opinions. Demonstrate that a claim is important or that a result is remarkable; then, leave readers to draw their own conclusions. Those conclusions will coincide with yours if you’ve demonstrated your point.
    4. Use the active voice, rather than the passive voice. Take responsibility for your work. Editors of high-impact scientific journals have endorsed this advice.
    5. Refer to yourself when necessary and only when necessary.
      • Example of unnecessary reference to self: We use the superscript “max” to signify the maximal Fisher information.
        Preferable alternative: The superscript “max” signifies the maximal Fisher information.
      • Example of unnecessary reference to self: Our results establish several opportunities for future research. First, we can implement the experimental proposals.
        Preferable alternative: Our results establish several opportunities for future research. First, one can implement the experimental proposals.
      • You may use the first-person plural when escorting the reader through a derivation.
        • Example: We substitute from Eq. (1) into Eq. (2).
    6. If you’re the only author, don’t use the plural (“we,” “our,” etc.). The usage is inaccurate and misleading. It portrays you as dodging responsibility for your work by dispersing that responsibility across the scientific community.
    7. Avoid dangling modifiers.
    8. Pair every verb with the appropriate noun.
      • Example of grammatically incorrect text: Equation (1) follows by calculating the sum.
        • One should pair the verb “calculate” with the noun “we,” because “we” undertook the calculating. However, this example’s author omitted the noun out of squeamishness about using the first person in a scientific document. Hence the sentence says that the equation calculates the sum. Equations can’t calculate sums.
      • Examples of correct alternatives
        • We derived Eq. (1) by calculating the sum.
        • Equation (1) follows from the evaluation of the sum.
        • Calculating the sum yields Eq. (1).
    9. Avoid empty subjects.
    10. Include no unnecessary words.
      1. “So-called” is unnecessary.
      2. “Note that” and “We note that” are unnecessary.
      3. Several phrases often preface mathematical statements but are unnecessary: “we have,” “we have that,” “it holds that,” and “it follows that.” One can better serve the reader by prefacing the mathematical statement with (i) a derivation or (ii) a prose description of the statement.
      4. Never write “is equal to”; “equals” is more concise.
      5. Never write “is able to”; “can” is more concise.
      6. Never write “gives an upper bound to” or “places an upper bound on”; “upper-bounds” is more concise. Analogous statements concern lower bounds.
      7. Never write “a large number of”; “many” is more concise. Never write “a small number of”; “few” is more concise.
      8. Never write “We refer to [symbol] as [name]”; “we call [symbol] [name]” is more concise.
      9. Never write “as long as”; “if” is more concise.
      10. The symbol > means “greater than”; and \geq, “greater than or equal to.” Don’t translate > into “strictly greater than”; the “strictly” is unnecessary. Analogous statements concern < and \leq.
    11. Avoid contractions, which are too informal for professional writing.
    12. The possessive is not a contraction and belongs in professional writing. It facilitates conciseness.
    13. Use the word “for” only when it belongs. Physicists often write “for” when they mean “if,” “per,” “at,” or something else.
      • Example of inappropriate use: The function vanishes for odd arguments.
        Corrected statement: If the argument is odd, the function vanishes.
      • Example of inappropriate use: We performed 10 trials for each parameter value.
        Corrected statement: We performed 10 trials per parameter value.
      • Example of inappropriate use: The function is smaller for small x values.
        Corrected statement: The function is smaller at small x values.
    14. Write “we evolve the state,” “we measure,” etc. only if you’re an experimentalist who undertakes those actions. Alternatives include “Consider measuring,” “Suppose the system evolves,” and the command tense (e.g., “One can measure this quantity as follows: prepare the qubit in \lvert 0\rangle. Evolve it under H…”).
    15. The condition x\ll y defines a regime, not a limit. The conditions \lim_{x\to0} and \lim_{y\to\infty} define limits and are inequivalent to x\ll y.
    16. Write “first,” “second,” “last,” etc., not “firstly,” “secondly,” “lastly,” etc. (I defer in this matter to The Elements of Style.)
    17. Humans can assume, suppose, etc. Mathematical expressions, protocols, etc. can’t.
    18. One multiplies factors together and sums terms. Don’t call factors terms and vice versa.
    19. If you mean “X equals Y,” say so. Don’t write “X agrees with Y,” “X matches Y,” or “we identify X with Y.” The latter three phrases are vaguer, and two of them contain more words, than “X equals Y.”
    20. Regarding the words “general” and “generally”:
      1. A general object subsumes every example of that object. If any example behaves unlike a supposedly general object, don’t call the object general.
      2. Many claims contain the term “general,” “generally,” or “in general” but don’t need the term.
        • Example of a sentence that contains “generally”: The terms generally commute.
        • Equivalent, more concise sentence: The terms commute.
      3. Physicists tend to use the words “general” and “generic” differently. By “general,” physicists usually mean “subsuming every example.” By “generic,” we usually mean “typical,” or “common.”
    21. “Then” makes sense (i) in discussions of chronology and (ii) in if–then statements. Don’t use “then” outside these contexts.
      • Example of inappropriate use: “Define X:=\ldots Then Y.”
      • Examples of appropriate alternatives
        • Define X:=\ldots This definition implies Y.
        • If X:=\ldots \, , then Y.
        • Define X:=\ldots \, , such that Y.
    22. Don’t justify any equation with “we used that [such-and-such is true],” which violates the rules of grammar. Grammatically correct alternatives include “We applied [a property],” “The equation follows from [a property],” and “…since [such-and-such is true].”
    23. Regarding tense:
      1. When describing what you’ve accomplished, use only one tense.
      2. Experiments happened in the past, so describe them in the past tense.
      3. When describing a proof’s steps, use the present tense.
        • Example: We Taylor-approximate the function about x=0. Substituting into Eq. (1) yields [equation].
    24. Nouns, verbs, and adjectives should agree about whether a quantity is singular or plural.
      • Example of singular–plural mismatch: The equations are a rule for evolving the cellular automaton.
      • Example alternative: The equations form a rule for evolving the cellular automaton.
    25. “Admit of” means “allow for,” or “permit.” The phrase needs the “of.”
      • Example: The formula admits of the following interpretation.
  3. Punctuation
    1. Consider any list that contains at least three items. If no item contains a comma, separate the items with commas. If any item contains a comma, separate the items with semicolons.
    2. In American English, periods and commas belong inside quotation marks. (Example: She told me, “Have a good day.”) In British English, periods and commas belong outside quotation marks. (Example: She told me, “Have a good day”.)
    3. To write quotation marks in LaTeX, don’t use your keyboard’s quotation-mark key; use the appropriate keys.
    4. Regarding hyphens:
      1. The hyphen (-) feeds into punctuation of three types: the hyphen (-), the en dash (–), and the em dash (—).
      2. The hyphen appears in some compound words, as in “non-negative.”
      3. In American English, em dashes can separate ideas within a sentence. Don’t separate any em dash from neighboring text with a space.
        • Example of appropriate use: The sample—the only product of this experiment—barely survived.
        • Example of inappropriate use: The sample — the only product of this experiment — barely survived.
        • Example of appropriate use: He told me only one sample had survived—hardly what I wanted to hear.
      4. This article specifies how to use the en dash. One use is “to separate the names of two or more people used as a compound modifier.”
        • Example: Feynman–Kitaev clock
      5. Hyphenate compound adjectives.
      6. If an adverb ends in “-ly,” it probably shouldn’t precede a hyphen.
        • Example of inappropriate hyphenation: strongly-coupled systems
      7. Follow a prefix with a hyphen if and only if the Physical Review style guide indicates that you should.
    5. A complete clause must follow any semicolon (unless the semicolon separates items in a list).
  4. Math
    1. Introduce only necessary notation, which readers will have enough trouble remembering. If a mathematical symbol appears only once, eliminate it. If a symbol appears only twice, try to eliminate it.
    2. Every sentence must obey the rules of English grammar, punctuation, and syntax, regardless of whether the sentence contains mathematical symbols. All math-containing sentences must end with punctuation marks. If a sentence contains a list of mathematical expressions, precede the final expressions with an “and.” If the list contains at least three mathematical expressions, separate them with commas.
    3. Introduce almost every mathematical symbol before you use it. If you introduce a symbol after using it, the reader will encounter the first use, stop, feel confused for a while, tentatively continue, find the definition, return to the earlier use to understand it, and then progress again. This back-and-forth breaks up the reading process. You may define a mathematical symbol after using it only if (i) the symbol is very common, known to nearly all physicists, and unmistakeable and (ii) defining the symbol earlier would disrupt the text’s flow.
    4. If you define a new function, denote it by only one letter. (I defer in this matter to the Physical Review style guide.)
      • Example: f(x,y,z)
      • Examples of disallowed notation: fxn(x,y,z), {\rm fxn}(x,y,z)
    5. Suppose that a superscript or subscript stands for a word or phrase without representing any variable or constant. The superscript/subscript must not be italicized. (I defer in this matter to Physical Review style guide.)
      • Example: Let x_{\mathrm{meas}} denote the measurement outcome.
    6. If a variable or constant appears in a superscript, parenthesize it. The parentheses communicate that the superscript isn’t an exponent.
      • Example: Let \sigma_z^{(j)} denote the Pauli-z operator of qubit j.
      • If a superscript is not italicized (stands for a word or phrase), don’t parenthesize it.
    7. Regarding the definition of a symbol A:
      1. If you write A alone on one side of a defining equation, use \coloneqq or \eqqcolon: A \coloneqq [expression], or [expression] \eqqcolon A. The symbols \coloneqq and \eqqcolon relate more information than does \equiv, encoding directionality.
      2. Use \equiv if A does not appear alone on its side of the equation: [function of A] \equiv [result of replacing A with its definition in the equation’s left-hand side].
    8. Refer to the Cartesian axes using the formatting “[italicized letter]-axis.” Don’t include any hat, boldface, or \vec symbol.
      • Example: x-axis
    9. Avoid denoting any index by i, which means \sqrt{-1} to physicists. Use j instead, unless you’re writing for engineers (who denote \sqrt{-1} by j).
    10. Don’t use the lowercase letter l (“ell”) as an index; readers might mistake it for a one. Use \ell (\ell) instead.
    11. Give every set-off equation a number. Readers (and coauthors) may want to refer to the equation easily when discussing the paper. Save them (and us) from having to say, e.g., “that equation halfway down page three.”
    12. When writing a set-off mathematical expression, use the align environment, not the equation environment. Using the align environment, one can easily extend an expression across multiple lines.
    13. Regarding a set-off mathematical expression that extends across multiple lines:
      1. Format the expression as follows by default.
        1. Put an & symbol immediately leftward of the first = sign or analogous symbol (e.g., \leq).
        2. If any subsequent line begins with another = sign (or analogous symbol), put an & immediately leftward of the symbol. (I’ll stop writing “or analogous symbol.”)
        3. Suppose that a subsequent line begins with a +, –, \times, or /. Find the symbol immediately rightward of the initial = sign. Begin the new line directly below that symbol.
        • Example:
      2. Modify the default formatting if necessary (a) to reduce the number of lines used in a PRL submission or (b) if the initial = appears awkwardly far to the right.
        • Example of (b):
      3. Suppose a new line begins with a term or factor, such as the jx^8 in the example under (A). Put the corresponding +, –, \times, or / at the beginning of the new line, not at the end of the previous line.
        • Examples of inappropriate placement:
    14. The symbol \approx means “approximately equals”; and ~, “scales as.” Approximations convey more information than scaling relations do.
    15. Use big-O-type notation or ~ symbols, not both; they’re partially redundant.
    16. \ldots, rather than \cdots, should stand in for elements that fit a pattern.
      • Example: x_1,x_2,\ldots,x_n
    17. When using \ldots as in the previous rule, present at least two initial examples of the pattern. One can’t define the pattern.
      • Contains insufficient examples: x_1,\ldots,x_n \, . For example, if n is odd, then x_1, x_2, \ldots, x_n and x_1, x_3, \ldots, x_n fit the template.
    18. Parentheses (), square brackets [], and curly braces {} are delimiters. If you nest them, do so in the order dictated by the Physical Review style guide.
    19. If delimiters enclose a symbol, it shouldn’t protrude above or below them (unless the delimiters would have to be grotesquely enormous). Use the \left and \right commands if the delimiters appear on the same line.
    20. An operator O isn’t a matrix; a matrix represents an operator in terms of a particular basis. Therefore, no equals sign should interrelate an O and a matrix. An arrow can.
      • Example: O\to\begin{bmatrix}1&0\\0&2\end{bmatrix}
    21. Every real number is complex. Don’t say “complex” if you mean “nonreal.”
    22. Consider introducing a mathematical symbol in a prose sentence without using a comma or colon. Put the symbol immediately after the word that names the object represented by the symbol.
      • Example of inappropriate placement: the set of real numbers \{ a, b \}
      • Examples of appropriate placements
        • the set \{a, b\} of real numbers
        • the set of real numbers a and b
        • Recall the set of real numbers, \{a, b\}, in Lemma 1.
  5. More mechanics of writing
    1. Use concise sentences, as advocated for in The Elements of Style. The reader can hold only so many ideas in their head at once.
    2. Structure sentences simply, as advocated for in The Elements of Style. The reader should be able to grasp each sentence easily.
      • Avoid nesting ideas within a sentence, to avoid convoluting the sentence’s structure.
        • Example of sentence with convoluted, nested structure: Any model of equilibrium and nonequilibrium behaviors of systems observed in tabletop experiments and high-energy colliders must obey the laws of relativistic quantum mechanics.
        • Visualization of the nesting: [Any model of ([(equilibrium and nonequilibrium) behaviors] of {systems observed in [(tabletop experiments) and (high-energy colliders)]})] must obey [the laws of (relativistic quantum mechanics)].
    3. The ideal paper title has the structure of a newspaper headline: it presents a claim, containing a subject and a predicate.
    4. Regarding abbreviations:
      1. Don’t abbreviate the first word in any sentence.
      2. Abbreviate “Figure,” “Section,” “Professor,” and “Appendix” if such a word appears partway through a sentence.
      3. Don’t abbreviate “Sections.”
    5. Regarding acronyms:
      1. Write every acronym in capital letters, as per the Physical Review style guide.
      2. Introduce each acronym the first time you use it.
      3. Thereafter, use only the acronym, not the spelled-out phrase, throughout the rest of the document’s main text. You may spell out the phrase in section, figure, and table titles if doing so improves the document’s clarity.
    6. Every paragraph should contain at least three sentences.
    7. Wherever you insert a blank line into your LateX code, a new paragraph begins in the corresponding PDF. Insert a blank line only if you wish to begin a new paragraph. This advice applies immediately before and after set-off equations.
    8. Never begin a subsection immediately after a section title. Between the two titles, overview the section. Analogous rules govern subsections and subsubsections.
    9. Put the word “only” in the appropriate place.
      • For example, suppose you’ve sampled data at a point x=0 in parameter space and sampled data at no other points. “We sampled data only at x=0” is correct; “We only sampled data at x=0” is probably not. The latter claim means that (i) you might have sampled data at x=0 and (ii) you did nothing to the x=0 data apart from sample it: you didn’t analyze the x=0 data, discuss the x=0 data, etc.
  6. When in doubt, consult the Physical Review style guide or The Elements of Style.
    • If those references don’t contain the information you seek, search for it in online writing guides. Not all such guides have equal merit, however. Lean toward guides written by human editors or published by college writing centers.

1 Collaborators have wondered why I use green; a student guessed it’s my favorite color. It isn’t; but I bleed green, having graduated from the Big Green, also known as Dartmouth College. Sometimes, I highlight certain pieces of text for one reason (e.g., to point out logical inconsistencies) and other text for another reason (e.g., to point out grammatical inconsistencies). Green distinguishes the first highlightings, while orange distinguishes the second: when not bleeding Dartmouth green, I bleed Caltech orange.

2 Don’t learn how to write from physics papers. 

September 26, 2026

Terence Tao — Cognitive Sanctuaries or: Manifesto of the Department Chair

[This is a guest post by Jess Werk. This blog post was initially written in a different file format and converted using AI. — T.]

Hello, I am Jess Werk, professor and chair of Astronomy at the University of Washington. Astronomy may interest mathematicians right now because our field has already been reshaped by supercomputers and survived to tell the tale. Our magneto-hydrodynamic simulations, built on the Navier-Stokes equations (plus lots of other “subgrid” physics), are run for over a hundred million core-hours, and produce emergent astrophysics that takes us years to understand and verify. They do not replace analytic theory; they complement it. The Rubin Observatory’s nightly stream of sky survey data, petabytes over the life of the project, would be unusable without advancements in cloud storage and AI algorithms. Astronomers were early adopters of machine-learning techniques because we have a whole Universe’s worth of beautiful data, and our space telescopes are built at the edge of what is technically feasible (e.g. the James Webb Space Telescope, and the Habitable Worlds Observatory now being designed). We are generally a technology-forward field, but a growing number of us see AI-enabled workflows as a risk to the profession. I will tell you what I think that risk is, and then put on my department chair hat and tell you what I am doing about it. My optimism is intact because it has to be.

As scientists, we design our experiments around an unattainable ideal: an objective and mindless observer from nowhere, studying a reality that exists separately from our subjective understanding of it. As creatures who study the cosmos from a relatively tiny rock, 93 million miles from the nearest star, embedded within the roiling gases of the Galactic interstellar medium produced by hundreds of billions of gasping stars, some fraction of which violently explode, we astronomers appreciate the physical impossibility of the observer from nowhere (and yet we endeavor to achieve it!). Thomas Nagel writes about the paradox of The View from Nowhere and argues later, in Mind and Cosmos, that an intelligible natural order which produces minds capable of understanding it is itself something that requires explanation. Intelligibility is the biological mind’s responsibility; every experiment we design must be understood by someone because that is what gives science its purpose. Generative AI attempts to embody the mindless observer from nowhere. Its achievements in both mathematics and science reveal the incompleteness of this long-held ideal and underscore the importance of human scientists and mathematicians, whose minds make results and proofs meaningful.

The marketing for and media coverage of generative AI invite a belief that the speed and volume of its output render slow human understanding meaningless. Thankfully, mathematicians have already begun to reject this idea. In astronomy, the result plays the role of the theorem: a paper is judged on what it found, how significant it was, and whether it found it first, far more than on what the finding means. The result serves as a proxy for understanding, and that proxy is now a problem (Kra, B., 2026). Across academia, incentive structures have long rewarded productivity over understanding, and generative AI now makes productivity easy to manufacture. Together, the two effects undermine the perceived value of doctoral-level study. Rather than swiftly implementing field-wide changes to our systems, I argue that protecting the Ph.D. requires simple, department-level fortifications that we can all make now. Together, these structural fortifications make up a model I have started calling the cognitive sanctuary.

Under the current productivity-weighted model, the technical debt that Henry Cohn describes in his blog post is borne by the most vulnerable members of our community. Ph.D. students compete for postdocs on the number of their publications; postdocs compete for faculty jobs on their h-indices; early-career faculty are judged on the dollar value of their grants (funders, in turn, count our “research products” which are ingested into a database) and the quantity of their scholarly output (including the number of Ph.D. students trained). These metrics do not measure or reward scientific merit. Citations accrue to papers that are already well-cited (e.g. Merton 1968, The Matthew Effect in Science), so the h-index inherits and magnifies biases present in the field (e.g. Kelly and Jennions 2006, Caplar, Tacchella and Birrer 2017). In the meantime, the literature keeps growing: astronomy arXiv submissions rose 14% from 2024 to 2025 (Lewis, Shah and Alfred 2026) and more than half of the papers posted in 2025 are written with language-model assistance (Saad and Ting 2026). Only about one paper in 66 that uses these AI tools declares it, a gap Saad and Ting attribute to failing disclosure norms and I would also attribute to perceived stigma. The people with the strongest incentive to use these tools and hide that they have done so are the same ones under the most pressure to publish: our students and early-career scientists. Unfortunately, they pay twice. Those who are still building new skills are the ones most likely to suffer “cognitive debt” from an overreliance on an LLM (e.g. Kosmyna et al. 2025; Bastani et al. 2025) and they will also face the steepest professional consequences for failing to declare its use (e.g., forced retractions and publication bans).

Scientific and mathematical thought relies on our collective judgment being tested, then re-tested and held up continually against alternative explanations. Our workflows, increasingly built on technology, have been streamlined to enable first discoveries and to decorate them with our own names. The slower, iterative processes of refining theories and interpreting data are relegated to methods sections that a few people might read, if they appear at all in the publication. A paper is rarely dedicated to showing how a particular idea turned out to be wrong after a careful analysis. “Nobody has time for that,” we think. The struggle of doing science is the struggle of learning, and it is also where the joy of discovery lives. A cognitive sanctuary, therefore, must reward the effort and celebrate a process that includes failure. It is a space that encourages unhurried thinking with plenty of room for rabbit holes and the occasional mad hatter.

Unlike a mad hatter, agentic AI creates a chain of probabilistic decisions and steers users away from the improbable. Some of my colleagues have envisioned a future in which student learning centers on individualized, guided conversations with an “Agentic Professor” (Cornillon and Prochaska 2026). In this future, the ability to derive and manipulate equations becomes secondary, and it is assumed that the capacity to interrogate a hypothesis, a necessary Ph.D.-level skill, can be built without the practice of trying and failing and trying and failing and trying and eventually succeeding. Sam Altman envisions a personal AI team for everyone and a virtual tutor for every child (Altman, S 2024). Taken to its end, the vision leaves no Ph.D. advisor to serve as a witness to the student’s judgment, and no one to vouch that the student can be trusted to produce and judge knowledge. I reject this future. The Ph.D. student cannot supervise mathematics without being able to produce it themselves through “active shaping of experience performed in the pursuit of knowledge” (Polanyi, M. 1966). Practice builds a tacit element of understanding that my field calls physical intuition. Nothing yet shows that physical intuition can be developed through closed-loop agentic AI conversations, and recent evidence points in the opposite direction (Bastani et al. 2025). Interrogation is best practiced among other scientists who can offer surprising alternatives and explanations (sometimes incorrect!), while an AI agent’s alternatives are drawn from the distribution of what has already been written. Our system of knowledge depends upon the Ph.D. and the Ph.D. depends upon iterative judgment of people who already have it.

And the people who have that judgment already work down the hall from you. Academic departments bring together kindred spirits and support many of the structures a cognitive sanctuary needs: we host seminars, discussions, hack-a-thons, and community events designed to honor thinking minds. What I see as a department chair is that the senior faculty with the most recognition are often the ones who use these structures least. Nobody wins an award for being a department chair, as I sadly have discovered, and the work of showing up to preprint discussions carries no service credit. When you are leading national committees and international collaborations, it is easy to justify skipping colloquium or a lunch talk for the hour gained in research productivity. Faculty are expected to do too much, and being overachievers, we do even more. Held against all that external and often unpaid labor, taking time to honor a half-formed thought process looks like an unaffordable luxury. Meanwhile, Ph.D. students are craving opportunities to interact with the senior faculty who are so often absent from department life. A department that works as a cognitive sanctuary gives senior faculty the cover to decline some external work, and it counts showing up for junior colleagues as critical service.

At our Astronomy faculty retreat on September 17, we discussed a scenario modeled on what is happening in mathematics. Briefly, OpenAI posts a preprint, press release, and public decision log of 40,000 steps reporting a five-sigma detection of an evolving dark energy equation of state, inconsistent with a cosmological constant, with error bars a factor of 2.5 tighter than anything our community could achieve from the same public data, drawn from a future survey in which our department has already invested heavily. None of the faculty were especially fazed and none thought the scenario to be implausible. Although they are all using AI in their research to different degrees, the faculty broadly agreed on four ideas:

  1. Generative AI is changing who has access to discovery, and early-career scientists face the steepest barriers.
  2. The interpretation of a discovery is more valuable than the discovery itself, and how we figure something out matters more than what we figure out.
  3. Scientific writing is best when it is slow and iterative because the writing refines the science (Gopen and Swan 1990).
  4. Verifying someone else’s work is less rewarding than doing your own, so a field that is about to depend on verification more than ever must reward it deliberately.

A cognitive sanctuary protects the thought process, for everyone in the department, but above all, for Ph.D. students. The Harvard Summit on PhD Math Education in the Age of AI met on the same day as our astronomy faculty retreat, and reached many of the same conclusions about assessment and AI use; it does not address faculty incentives, which is where much of the department’s leverage lies. Below are several practical suggestions for how to build a cognitive sanctuary in your own department. They fall into three groups: Ph.D. processes, faculty reward systems, and department community.

Ph.D. Processes

  1. Develop detailed rubrics for both the dissertation and the defense that distinguish a pass from a conditional pass from a failure. One dimension, for example, is whether the student can state the competing interpretations of the result and say what measurement would distinguish them. Require committee members to score the rubric independently before any closed-door discussion. Following a successful defense, require the committee to sign a short report stating what the student demonstrated. The report can go into recommendation letters, and it gives hiring committees something other than a publication list to read. Post the rubric and the evaluation process on the department website.
  2. Design any evaluation checkpoints (e.g. qualifying exams, required committee meetings) to be conducted live, preferably in person, and unassisted.
  3. Remove any requirement that a student publish a paper to advance to candidacy. Do not push students to submit publications early, before interpretations have been discussed and worked out. Encourage faculty and Ph.D. students to submit short write-ups of ideas that did not pan out.
  4. Ask Ph.D. candidates to explain their work at the board, in group meetings and once a year to the whole department. Explaining an idea to a room of scientists helps refine it, and explaining how an idea turned out to be wrong is part of learning. Ask faculty to provide feedback on the substance of the explanation, not the delivery.
  5. Set clear boundaries on AI use in Ph.D. research to honor the learning process. These boundaries apply to advisors and external collaborators as well as to students. Agree that advisors will not send students feedback that is written with generative AI. Agree that students will not send advisors text or analysis generated by AI.
  6. Teach students to interrogate AI model outputs on well-understood problems and build these skills in coursework, and group and one-on-one meetings.

Faculty Reward Systems

  1. Define unit-level promotion and tenure criteria that consider Ph.D. student mentoring to be at least as important as a prestigious award or large grant.
  2. Publicly post faculty department service assignments, and weight Ph.D. student advising and committee work more heavily than external service (including university-level service). Provide service credits for faculty who commit to regularly show up to key department events (e.g. chalkboard talks, colloquium).
  3. Develop a teaching policy that includes credit for Ph.D. student advising and committee work. For example, a faculty member who sits on three or more Ph.D. committees in each of two consecutive years is offered a quarter of teaching relief once every 2 years.
  4. Set an expected external service load for faculty that they may cite when they decline invitations. Encourage faculty to report declined external service alongside accepted service in their annual merit reports.

Department Community

  1. Create opportunities for meaningful connection among senior faculty and students, e.g. department coffee hours, potlucks, celebrations of student milestones.
  2. Thoughtfully design colloquium and seminars with time for discussion, including dedicated time for Ph.D. students to meet with speakers. Commit to showing up yourself, even when speakers are discussing a topic outside of your area of expertise.
  3. Set a department-wide standard that no one uses AI-generated text in research publications, except for disclosed, light use for editing, grammar, or translation.
  4. Develop department-specific policies on generative AI use and have regular discussions with faculty on whether these policies are achieving their stated purpose.

None of these suggested structural fortifications presents a case against generative AI, a remarkable tool that can improve the practice of science and mathematics and reduce the burden of some administrative tasks. Cognitive sanctuaries encourage its use in the open, especially where nothing the student is supposed to be learning is at stake. The suggestions above do not address structural problems in our fields that will be amplified by AI. They will not themselves generate much-needed funding for basic science and mathematics at the national level. They cannot solve emerging issues in grade-school education that increasingly relies on AI.

Departments, and the Ph.D. students they educate, sit where all the above challenges meet. They absorb the funding cuts and receive the students that grade-school education produces. As we create microcosms of joy and learning in our departments, I hope that we can gradually steer universities toward rewarding process. Cognitive sanctuaries build the minds that make discoveries mean something, the “deployable intellectual reserve” that Amit Sahai envisions. In an optimistic future where AI-driven discoveries and innovations must be understood and tested, academia is intentionally restructured around process, collaborations that demonstrate understanding win the highest awards, and Ph.D. students find joy as their engaged human professors challenge them over and over again until what they know expands, and they can vouch for all of it.

Acknowledgements

Thank you to Tatiana Toro for the introduction to Terry, and to Terry for hosting this piece. Thank you also to Terry for pointing me to the report of the Harvard Summit, which I read only after finishing a full draft of this piece. I was genuinely heartened to find how much our two fields had converged on their own. Thank you to the faculty in the UW Department of Astronomy who inspire me with their thoughtfulness, brilliance and musical talent. Thank you to Professor Xavier Prochaska, my postdoc mentor who frequently had me go to his chalkboard with my ideas. He came to me over a year ago with a claim that, under his guidance, Claude could write a Ph.D. thesis in a month that would be better than that of an average Ph.D. student in Astronomy. I did not believe him at the time. I do now, but I have decided that the thesis is not the point of the Ph.D. The student is.

I found a strange comfort in reading philosophy this summer, particularly The View from Nowhere and Mind and Cosmos, both by Thomas Nagel, after my guitar teacher was killed in an accident on June 1st. He taught me how repetition builds skill, always turned on the metronome, and taught me intricate finger patterns. One day I would be all tangled up in a finger pattern, and the next it would fall into place. Our lessons were a microcosm of joy, a cognitive sanctuary, while I was suffering from the burnout of my first years as department chair. Our time together meant more to me than I can say, and what I learned from him will stay with me forever.

AI disclosure: I drafted every paragraph of this piece by hand, then typed and edited it. I used Claude (Fable 5.1) to check my sentences against Gopen and Swan 1990, to find and verify citations, and to argue with. A handful of sentences began as its suggested wording and were rewritten by me; the arguments are mine, though some were sharpened in the arguing. No text in this piece was generated and pasted.

References

Altman, S. 2024. “The Intelligence Age.” Blog post, September 23, 2024. https://ia.samaltman.com/

American Astronomical Society. 2026. “Author Guidelines for Use of AI and LLMs in Manuscript Preparation.” AAS Journals, posted September 9, 2026. https://journals.aas.org/author-llm-guidelines

Avila, A., et al. 2026. “A Severe Misalignment of AI in Mathematics.” Zenodo, September 11, 2026. https://doi.org/10.5281/zenodo.22737751

Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., and Mariman, R. 2025. “Generative AI without guardrails can harm learning: Evidence from high school mathematics.” Proceedings of the National Academy of Sciences 122 (26): e2422633122. https://doi.org/10.1073/pnas.2422633122

Caplar, N., Tacchella, S., and Birrer, S. 2017. “Quantitative evaluation of gender bias in astronomical publications from citation counts.” Nature Astronomy 1: 0141. https://doi.org/10.1038/s41550-017-0141

Cohn, H. 2026. “The technical debt of AI-generated mathematics.” Guest post, What’s new (Terence Tao’s blog), September 15, 2026. https://terrytao.wordpress.com/2026/09/15/the-technical-debt-of-ai-generated-mathematics/

Cornillon, P., and Prochaska, J. X. 2026. “The Agentic Professor: Exploring GenAI-Supported Futures in Higher Education.” EDUCAUSE Review, September 14, 2026. https://er.educause.edu/articles/2026/9/the-agentic-professor-exploring-genai-supported-futures-in-higher-education

Gopen, G. D., and Swan, J. A. 1990. “The Science of Scientific Writing.” American Scientist 78 (6): 550–558.

Kelly, C. D., and Jennions, M. D. 2006. “The h index and career assessment by numbers.” Trends in Ecology & Evolution 21 (4): 167–170. https://doi.org/10.1016/j.tree.2006.01.005

Kosmyna, N., et al. 2025. “Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task.” arXiv:2506.08872. https://arxiv.org/abs/2506.08872

Kra, B. 2026. “Deep theorems were scarce and difficult and so became an effective mechanism to identify deep thought. AI has broken this system.” Guest post, What’s new (Terence Tao’s blog), September 13, 2026. https://terrytao.wordpress.com/2026/09/13/deep-theorems-were-scarce-and-difficult-and-so-became-an-effective-mechanism-to-identify-deep-thought-ai-has-broken-this-system/

Lewis, R., Shah, H., and Alfred, A. 2026. “Astrophysics Wrapped 2025.” arXiv:2602.12303. https://arxiv.org/abs/2602.12303

Merton, R. K. 1968. “The Matthew Effect in Science.” Science 159 (3810): 56–63. https://doi.org/10.1126/science.159.3810.56

Nagel, T. 1986. The View from Nowhere. New York: Oxford University Press.

Nagel, T. 2012. Mind and Cosmos: Why the Materialist Neo-Darwinian Conception of Nature Is Almost Certainly False. New York: Oxford University Press.

Polanyi, M. 1966. The Tacit Dimension. Garden City, NY: Doubleday. Reissued 2009, Chicago: University of Chicago Press.

Saad, S. M., and Ting, Y.-S. 2026. “More than half of recent astronomy papers are written with language-model assistance.” arXiv:2609.10664. https://arxiv.org/abs/2609.10664

Sahai, A. 2026. “We’re gonna need a lot more mathematicians.” Guest post, What’s new (Terence Tao’s blog), September 24, 2026. https://terrytao.wordpress.com/2026/09/24/were-gonna-need-a-lot-more-mathematicians/

Sanderson, G. 2026. “If math is more than proof, we need to better celebrate the rest of it.” Guest post, What’s new (Terence Tao’s blog), September 18, 2026. https://terrytao.wordpress.com/2026/09/18/if-math-is-more-than-proof-we-need-to-better-celebrate-the-rest-of-it/

Summit on PhD Math Education in the Age of AI. 2026. Report of Summit on PhD Math Education in the Age of AI, September 17–18, 2026. Harvard Center of Mathematical Sciences and Applications. https://cmsa.fas.harvard.edu/media/2026/09/Summit-on-PhD-Math-Education-in-the-Age-of-AI.pdf

September 25, 2026

Jordan Ellenberg — I love working here, II

Making final revisions to Don’t Be Too Sure on the Terrace. An all-tuba band on the Terrace Stage playing an all-tuba cover of “Black Hole Sun.”

Sorry for light blogging lately. See “making final revisions” above. Just realized I misspelled the first name of Katharine Briggs, who’s mentioned several times and who I read a whole book about. How many mistakes are left in this thing? (I suppose if I’d kept a rigorous log of mistakes found and plotted it over time, I could curve-fit and estimate this.)

n-Category Café Binomial Coefficient Coincidences

These seven equations between binomial coefficients are ‘coincidences’ — they aren’t among the four known infinite families of such equations:

(162)=(103)=120 \binom{16}{2} = \binom{10}{3} = 120

(212)=(104)=210 \binom{21}{2} = \binom{10}{4} = 210

(562)=(223)=1540 \binom{56}{2} = \binom{22}{3} = 1540

(782)=(146)=3003 \binom{78}{2} = \binom{14}{6} = 3003

(1202)=(363)=7140 \binom{120}{2} = \binom{36}{3} = 7140

(1532)=(195)=11628 \binom{153}{2} = \binom{19}{5} = 11628

(2212)=(178)=24310 \binom{221}{2} = \binom{17}{8} = 24310

In 1997, De Weger conjectured that every equation between binomial coefficients follows from the four known systematic equations and the seven coincidences I showed you:

• Benjamin M. M. de Weger, Equal binomial coefficients: some elementary considerations, Journal of Number Theory 63, no. 2 (1997), 373–386.

At that time, he and his collaborators checked there were no others involving binomial coefficients less than 1030. Later they checked that there are none involving binomial coefficients less than 1060:

• Aart Blokhuis, Andries Brouwer and Benne de Weger, Binomial collisions and near collisions.

So, De Weger’s conjecture stands open. The four infinite families, by the way, are these:

(nk) = (nn−k) 0≤k≤n (n0) = 1 n≥0 ((nk)1) = (nk) 0≤k≤n \begin{array}{cccl} \displaystyle{ \binom{n}{k} } &=& \displaystyle{ \binom{n}{n-k} }& \qquad 0 \le k \le n \\ \\ \displaystyle{ \binom{n}{0} } &=& 1 & \qquad n \ge 0 \\ \\ \displaystyle{ \binom{\binom{n}{k}}{1} } &=& \displaystyle{\binom{n}{k}} & \qquad 0 \le k \le n \end{array}

and the only nontrivial one: the Lind–Singmaster family involving the Fibonacci numbers F iF_i where F 0=0,F 1=1F_0 = 0,\ F_1 = 1:

(F 2i+2F 2i+3F 2iF 2i+3)=(F 2i+2F 2i+3−1F 2iF 2i+3+1)i=1,2,3,… \displaystyle{ \binom{F_{2i+2}F_{2i+3}}{\,F_{2i}F_{2i+3}\,} \;=\; \binom{F_{2i+2}F_{2i+3}-1}{\,F_{2i}F_{2i+3}+1} \qquad i = 1,2,3,\dots }

The first three equations in the Lind–Singmaster family are these:

(155) = (146) (10439) = (10340) (714272) = (713273) \begin{array}{ccc} \displaystyle{ \binom{15}{5} } &amp;= &amp;\displaystyle{\binom{14}{6}} \\ \\ \displaystyle{\binom{104}{39} } &amp;= &amp;\displaystyle{\binom{103}{40}} \\ \\ \displaystyle{ \binom{714}{272} } &amp;=&amp; \displaystyle{\binom{713}{273}} \end{array}

When I said De Weger conjectured every equation between binomial coefficients follows from the four known systematic equations and the seven coincidences, I really meant it. For example, the first of the Lind-Singmaster equations combines with the coincidence

(782)=(146)=3003 \binom{78}{2} = \binom{14}{6} = 3003

and other systematic equations to give

(30031) = (782) = (155) = (146) = (30033002) = (7876) = (1510) = (148) \begin{array}{cccccccccc} & \binom{3003}{1} &=& \binom{78}{2} &=& \binom{15}{5} &=& \binom{14}{6} \\ \\ = & \binom{3003}{3002} &=& \binom{78}{76} &=& \binom{15}{10} &=& \binom{14}{8} \end{array}

In fact, if de Weger’s conjecture is true, the number 3003 shows up as a binomial coefficient in more ways than any other positive integer!

I’ll explain the Lind–Singmaster family later. Here’s the question I’m most interested in now:

Is there any good explanation for the seven binomial coefficient coincidences?

My collaborator Paul Schwahn found a beautiful explanation of the first one, namely

(103)=(162) \displaystyle{ \binom{10}{3} = \binom{16}{2} }

His explanation uses representation theory. The Lie algebra 𝔰𝔬(10)\mathfrak{so}(10) has a 10-dimensional representation, the ‘vector’ representation V 10V_{10}, and also two 16-dimensional representations, the ‘left and right-handed spinor’ representations S 10 ±.S^\pm_{10}. There’s an isomorphism of representations

Λ 2S 10 +≅Λ 3V 10 \displaystyle{\Lambda^2 S^+_{10} \cong \Lambda^3 V_{10} }

and similarly for S 10 −,S^-_{10}, but we might as well work with S 10 +.S^+_{10}. Here Λ k\Lambda^k means the kth exterior power. For any vector space XX we have

dim(Λ kX)=(dimXk) \displaystyle{ \dim(\Lambda^k X) = \binom{\dim X}{k} }

Thus, taking dimensions, the isomorphism of representations

Λ 2S 10 +≅Λ 3V 10 \displaystyle{\Lambda^2 S^+_{10} \cong \Lambda^3 V_{10} }

instantly gives

(162)=(103) \displaystyle{ \binom{16}{2} = \binom{10}{3} }

It is not super-easy to prove this isomorphism of representations, but it’s still nice to find a deeper layer of meaning underlying what might otherwise seem like a meaningless coincidence!

Can we find good explanations for the other six coincidences?

First let me explain two failed attempts using representation theory, and then three successes using combinatorics.

Representation theory

First: we can look for isomorphisms like

Λ 2S 10 +≅Λ 3V 10 \displaystyle{\Lambda^2 S^+_{10} \cong \Lambda^3 V_{10} }

involving representations of 𝔰𝔬(n)\mathfrak{so}(n) for larger n.n. In fact this isomorphism is part of a pattern! The next one involves the vector and left-handed spinor representations of 𝔰𝔬(12).\mathfrak{so}(12). But it’s this:

Λ 2S 12 +≅Λ 4V 12⊕Λ 0V 12 \displaystyle{ \Lambda^2 S^+_{12} \cong \Lambda^4 V_{12} \oplus \Lambda^0 V_{12} }

so it gives

(322)=(124)+1 \displaystyle{ \binom{32}{2} = \binom{12}{4} + 1 }

or

496=495+1 \displaystyle{ 496 = 495 + 1 }

So we fail to get an equation between binomial coefficients, but we’re close: we’re only off by one.

The second paper I cited, Binomial collisions and near collisions, presents a list of cases where two binomial coefficients differ by one. This is on the list. So we failed to explain an equation between binomial coefficients, but we explained a near-miss.

Here’s another failed attempt at explaining an equation between binomial coefficients. The equation

(782)=(146)=3003 \displaystyle{ \binom{78}{2} = \binom{14}{6} = 3003 }

is fascinating to anyone who knows their exceptional Lie groups. 78 is the dimension of E 6,\mathrm{E}_6, while 14 is the dimension of G 2.\mathrm{G}_2. G 2\mathrm{G}_2 is a subgroup of E 6\mathrm{E}_6 because G 2\mathrm{G}_2 is the automorphism group of the octonions and E 6\mathrm{E}_6 is the isometry group of the bioctonionic plane. We’d get the above equation if the 2nd exterior power of the adjoint representation of E 6,\mathrm{E}_6, upon being restricted to G 2,\mathrm{G}_2, were isomorphic to the 6th exterior power of the adjoint representation of G 2.\mathrm{G}_2.

Amazingly, it seems these two representations of G 2\mathrm{G}_2 are not isomorphic even though their dimensions are the same: both 3003.

Even more amazingly, E 6\mathrm{E}_6 and G 2\mathrm{G}_2 both have irreducible representations of dimension 3003, but they are not the representations I just mentioned.

I would be happy for someone to check these two claims.

Combinatorics

It turns out that the following three coincidences can all be explained by the same style of combinatorial argument:

(212)=(104),(1532)=(195),(782)=(146) \binom{21}{2} = \binom{10}{4}, \qquad \binom{153}{2} = \binom{19}{5}, \qquad \binom{78}{2} = \binom{14}{6}

The argument is not very elegant, but it’s moderately interesting. Mike Stay got it started by asking ChatGPT, and I finished it off.

In each case the argument has three steps. The first two steps are what combinatorialists call bijective proofs: we prove an equation between numbers by proving a natural isomorphism between structures on sets and then taking cardinalities. The third step is non-bijective because it involves taking an equation and dividing both sides by the same number. Maybe we can make it closer to bijective by using groupoid cardinality, which allows for division, but I haven’t tried that.

Step 1: pairs of edges in a complete graph

Starting from the left-hand binomial coefficient in each equation we’re trying to prove, note that the number on top is a triangular number:

21=(72),153=(182),78=(132) 21 = \binom{7}{2}, \qquad 153 = \binom{18}{2}, \qquad 78 = \binom{13}{2}

Note that ((n2)2)\binom{\binom{n}{2}}{2} counts unordered pairs of distinct edges in the complete graph on nn vertices. Two distinct edges either share a vertex or are disjoint, so there are two cases:

• Sharing a vertex: the pair spans 3 vertices and is determined by this 3-element set together with a choice of which vertex is shared. That gives 3(n3)3\binom{n}{3} pairs.

• Disjoint: the pair spans 4 vertices and is determined by this 4-element set together with one of its 3 splittings into two pairs. That gives 3(n4)3 \binom{n}{4} pairs.

By Pascal’s rule, (n3)+(n4)=(n+14).\binom{n}{3} + \binom{n}{4} = \binom{n+1}{4}. But this also has a bijective proof: add a new point to the set of nn vertices, and a 4-element subset of the enlarged set either contains the new point (so involves a 3-element subset of the old ones) or does not (so involves a 4-element subset of the old ones). Hence we have a bijective proof that

((n2)2)=3(n+14) \binom{\binom{n}{2}}{2} = 3\binom{n+1}{4}

This holds for all n.n. For our three cases we get bijective proofs of the following equations:

(212)=3(84),(1532)=3(194),(782)=3(144). \binom{21}{2} = 3\binom{8}{4}, \qquad \binom{153}{2} = 3\binom{19}{4}, \qquad \binom{78}{2} = 3\binom{14}{4}.

It thus remains to prove

3(84)=(104),3(194)=(195),3(144)=(146). 3\binom{8}{4} = \binom{10}{4}, \qquad 3\binom{19}{4} = \binom{19}{5}, \qquad 3\binom{14}{4} = \binom{14}{6}.

Step 2: a double count

Each of the above equations follows by counting a single set in two ways, and then a nonbijective step: dividing these counts by the same number.

• Showing 3(84)=(104)3\binom{8}{4} = \binom{10}{4}. In a 10-element set, count triples (S,x,y)(S, x, y) in two different ways, where SS is a 4-element subset and x,yx, y are distinct points outside SS. Choosing SS first we see there are (104)⋅6⋅5=30(104) \binom{10}{4} \cdot 6 \cdot 5 = 30\binom{10}{4} triples. Choosing xx and yy first we see there are 10⋅9⋅(84)=90(84) 10 \cdot 9 \cdot \binom{8}{4} = 90\binom{8}{4} triples. So, we get a bijective proof that 90(84)=30(104) 90\binom{8}{4} = 30\binom{10}{4} You can see what we’ll do next.

• Showing 3(194)=(195).3\binom{19}{4} = \binom{19}{5}. In a 19-element set, count pairs A⊂BA \subset B with |A|=4|A| = 4 and |B|=5|B| = 5. On the one hand, there are (194)\binom{19}{4} choices of A,A, and for each there are 15 choices of BB since we can add an extra point in 15 ways. On the other hand, there are (195)\binom{19}{5} choices of B,B, and for each there are 5 choices of AA since we can choose AA in (54)=5\binom{5}{4} = 5 ways. So we get a bijective proof that 15(194)=5(195) 15\binom{19}{4} = 5\binom{19}{5} Again, you can see what we’ll do next!

• Showing 3(144)=(146).3\binom{14}{4} = \binom{14}{6}. In a 14-element set, count pairs A⊂BA \subset B with |A|=4|A| = 4 and |B|=6|B| = 6. On the one hand, there are (144)\binom{14}{4} choices of A,A, and for each there are 45 choices of BB since we can add 2 other points in (102)=45\binom{10}{2} = 45 ways. On the other hand there are (146)\binom{14}{6} choices of BB and for each there are 15 choices of AA since we can choose 4 points in (64)=15\binom{6}{4} = 15 ways. This gives a bijective proof that 45(144)=15(146) 45\binom{14}{4} = 15\binom{14}{6} Again you can see what we’ll do next.

Step 3: division

Having proved

90(84)=30(104),15(194)=5(195),45(144)=15(146) 90\binom{8}{4} = 30\binom{10}{4}, \qquad 15\binom{19}{4} = 5\binom{19}{5}, \quad 45\binom{14}{4} = 15\binom{14}{6}

we can now divide by 30, 5 and 15, respectively, and get

3(84)=(104),3(194)=(195),3(144)=(146) 3 \binom{8}{4} = \binom{10}{4} , \qquad 3 \binom{19}{4} = \binom{19}{5}, \quad 3\binom{14}{4} = \binom{14}{6}

as we wanted, completing our proof that

(212)=(104),(1532)=(195),(782)=(146). \binom{21}{2} = \binom{10}{4}, \qquad \binom{153}{2} = \binom{19}{5}, \qquad \binom{78}{2} = \binom{14}{6}.

I don’t see how to use tricks of the same general sort to explain the remaining coincidences

(562)=(223),(1202)=(363),(2212)=(178). \binom{56}{2} = \binom{22}{3}, \qquad \binom{120}{2} = \binom{36}{3}, \qquad \binom{221}{2} = \binom{17}{8}.

The Lind–Singmaster family

Lind and Singmaster were trying to find all n,kn,k with

(nk)=(n−1k+1) \displaystyle{ \binom{n}{k} = \binom{n-1}{k+1} }

I’ll rapidly sketch the key steps of their argument. Simplifying the equation above we get

n(k+1)=(n−k)(n−k−1) n(k+1) = (n-k)(n-k-1)

or

n 2−(3k+2)n+(k 2+k)=0 n^2 - (3k+2)n + (k^2+k) = 0

Solve for nn using the quadratic formula. This formula turns out to have

5k 2+8k+4 \sqrt{5k^2 +8k+4}

in it. So we need 5k 2+8k+45k^2 +8k+4 to be a perfect square!

Now we’re trying to find integer solutions of

5k 2+8k+4=m 2 5k^2+8k+4 = m^2

A quadratic diophantine equation! Multiply by 5 and complete the square:

5m 2=(5k+4) 2+4 5m^2 = (5k+4)^2 + 4

y=5k+4y = 5k+4 is an integer when kk is, so we need to find integer solutions of

y 2−5m 2=−4 y^2 - 5m^2 = -4

This is a ‘Pell equation’, and people know how to solve these. In this particular case we get all the solutions from this fact:

L n 2−5F n 2=4(−1) n L_n^2 - 5F_n^2 = 4(-1)^n

where F nF_n are the Fibonacci numbers 0, 1, 1, 2, 3, … and L nL_n are the Lucas numbers 2, 1, 3, 4, 7, …. These are two sequences satisfying the same famous recurrence relation, just with different initial conditions.

We want nn odd, to get

L n 2−5F n 2=−4 L_n^2 - 5F_n^2 = -4

It turns out y=L n,m=F ny = L_n, m = F_n with nn odd give all solutions of the Pell equation

y 2−5m 2=−4 y^2 - 5m^2 = -4

However, remember I said y=5k+4y = 5k+4 is an integer when kk is. But the converse isn’t always true, and we need kk to be an integer! This clearly happens iff y≡4bmod5.y \equiv 4 \bmod 5.

So we need to know when L n≡4bmod5L_n \equiv 4 \bmod 5 Apparently this happens iff n≡3bmod4.n \equiv 3 \bmod 4. I won’t think about this now… but this is the last hard step.

In summary, we’ve seen

(nk)=(n−1k+1) \displaystyle{ \binom{n}{k} = \binom{n-1}{k+1} }

if and only if y=5k+4y = 5k+4 is a Lucas number L nL_n with n≡3bmod4.n \equiv 3 \bmod 4. We could quit here, but people like to use the identity

L 4i+3−4=5F 2iF 2i+3 L_{4i+3} - 4 = 5F_{2i} F_{2i+3}

to get a formula for kk in terms of Fibonacci numbers. This is gilding the lily, I’d say, but it eventually leads to the formula that de Weger presents:

(F 2i+2F 2i+3F 2iF 2i+3)=(F 2i+2F 2i+3−1F 2iF 2i+3+1),i=1,2,3,… \displaystyle{ \binom{F_{2i+2}F_{2i+3}}{\,F_{2i}F_{2i+3}\,} \;=\; \binom{F_{2i+2}F_{2i+3}-1}{\,F_{2i}F_{2i+3}+1\,}, \qquad i = 1,2,3,\dots }

The takeaway message is: our problem can easily be reduced to a quadratic diophantine equation, then put in Pell form… and it’s known that the sequence of integer solutions of a Pell equation obeys a linear recurrence relation! We luck out in this case and get solutions connected to Lucas and Fibonacci numbers.

It’s probably a lot easier to show that the Lind–Singmaster equations hold. The above argument does a lot more, by showing that every equation of the form

(nk)=(n−1k+1) \displaystyle{ \binom{n}{k} = \binom{n-1}{k+1} }

is in the Lind–Singmaster family. De Weger gets into a lot more number theory trying to rule out various other kinds of coincidences between binomial coefficients. It’s downright scary how these questions pull you into the deep, icy waters of mathematics.

Matt von Hippel — It Got to My Field

For the folks who found this blog from Anthropic’s site, welcome!

For everyone else, I’d better give some context.

Last month, I had a blog post titled “It Only Counts When AI Gets to My Field”. The title was a joke, the content less so. I said that if some of the longstanding problems of my old research field got solved by AI, then I’d sit up and take notice.

The post turned out to be a bit of a self-fulfilling prophecy, after some folks at Anthropic read it and decided to tackle one of the outstanding problems I mentioned. They then invited me to write a post about it on their science blog.

The post is here. I recommend reading it, then coming back. The rest of this post will be a little Q&A.

Q: At the end of that post, it says Anthropic paid you for your time. Can we trust what you wrote?

A: When I agreed to write the post, I made it clear that I was going to give my own opinion, not write an ad for Anthropic. They could propose light edits, but that’s it. And I avoided signing anything with them, not even an NDA, so I could freely tell you if they pushed my boundaries.

They didn’t push my boundaries. They asked me to clarify a few things, and to give more detail on the science. I didn’t change the message or the takeaways.

The reason I asked them to pay me is that, as a freelance writer, I don’t have a salary to fall back on. Time I spend on a project like that post is time I’m not spending on journalistic projects, so if I worked on it for free I’d essentially be using my vacation time for it. The rate I’m charging them is roughly in the middle of what I’d have gotten paid if I used that time on journalism: a bit more than I would have made working for the lower-paying outlets, a bit less than I would have made with the higher-paying ones.

I’m also not expecting this to be the start of a longer-term business relationship or anything like that. So overall, I don’t think I’m incentivized to lie on their behalf. You can trust me.

Q: How confident are you that they did what they said they did?

A: I didn’t get the feeling they were lying. But I didn’t start out skeptical.

I haven’t seen the logs from the LLM, or anything like that. That’s the kind of thing I would dig into if I were more suspicious of their story. If there’s a reason to be suspicious, I’ll ask.

But so far, I don’t feel that I have much reason to be suspicious. They were able to look up details on the fly when I asked them, they didn’t seem to have a polished message they were trying to push past me. And more importantly, as I’ll mention in a bit, I don’t think anything they described is all that outlandish. It lines up, largely, with the capabilities I’d expect their Claude Science platform to have.

It’s also relevant that they didn’t use an internal model for this. It means that scientists will likely be trying out similar problems soon, so if it does turn out they exaggerated something, people are going to figure out quite quickly.

Q: So how big is this? We just saw AI solve a Millennium problem, after all.

A: This was a lot easier than a Millennium problem. But it was also a lot cheaper.

The problem they solved is one with a clear recipe, honed and explained over multiple papers. It’s something I expected to be hard to do without access to a lot of computer power, so I thought it would need to be approached with a novel technique, in order to avoid using that computer power.

In the end, it didn’t need that. As I mention in the post, another group got most of the result at around the same time, with a much smaller amount of AI assistance.

It’s not an easy recipe, to be clear. I think other people will be surprised that a reasonably affordable program like Claude Science can do this without a lot of guidance. I’m not surprised, mostly, because I’ve been paying enough attention to what people have been doing with these things, and carrying out this kind of recipe consistently is actually something the good science AIs can do right now. If you haven’t been following as closely, that’s going to be a lot more surprising.

So maybe the best way to answer this is: instead of paying $15 million to solve one of the most famous problems in mathematics, they paid $1,000 or so to make a significant next step in an ongoing research program in theoretical physics. This doesn’t tell you much about whether AI can achieve field-defining breakthroughs, but it tells you a lot about the kinds of things a theoretical physicist with $1000 spare budget can do right now.

(As an aside, it feels crazy to me that 90% or more of the compute cost was from the LLM, not the calculation itself. On the one hand, it feels nuts to essentially use ten times more computer power to do this than would have been needed if a human had done the coding. On the other hand, I’ve applied for grants that budgeted more than that per year for travel costs, and spending this kind of money on getting a result seems a lot more useful than spending it on airfare.)

Q: Isn’t it reckless to publicly post a problem for AI like that?

A: I posted my challenge before OpenAI announced their Navier-Stokes result. At that point, there had been a few awkward surprises, but for the most part AI companies had a pattern of talking to an academic first before trying to solve a problem. It’s what they did back in March.

I expected that was what they’d do this time, and got more than a little blindsided when they just solved the problem on their own and reached out to Lance afterwards. I’m lucky that Lance doesn’t seem mad at me about it, it would have been quite understandable if he was.

I did tell them that if they wanted to tackle any of the other problems in that post, they should reach out to one of the scientists involved first, before attempting it. In addition to being polite, it’s a way to make sure that they understand the problem correctly and have a plan to verify the result. They happened to pick a challenge that was particularly easy to verify, but it didn’t have to go that way.

Q: Any thoughts about AI’s future impact on science?

A: If there’s a problem you can’t solve, but you think someone with expertise in another field or better programming skills could, then it’s probably solvable with AI.

A lot of problems don’t fall into that category. I’ve seen people talking online about AI solving quantum gravity, but solving quantum gravity is a question of which bullets you’re willing to bite, not a question of technical skill. That doesn’t mean AI will never be able to address it, but if so it will come from some sort of superpersuader aspect, not merely scientific capabilities.

And of course, some fields do require experiments. People are increasingly building AI-powered labs. I’ll leave it to people from actual experimental fields to think about the potential there.

Q: Can you say something explicit about the bigger risks from AI, now?

A: As I mention in the post, I didn’t learn that much about the bigger questions from this. I’m still not an expert.

But there’s one thing I think is worth emphasizing:

This technology is clearly getting more effective over time. I don’t think they could have done this six months ago. If you’re trying to predict what will happen next, you shouldn’t just assume this is the most powerful it will get, or the cheapest. If you think there’s a limit, you need to argue for it.

Edit: One more thing I should mention, now that I’ve confirmed she’s ok with it: the idea for my challenge came from some discussions at Lancefest, a conference in honor of Lance back in June, and particularly from Anastasia Volovich, whose talk declared the nine-loop calculation “a benchmark – or a ‘challenge’ – against which any “AI takeover” should be measured”.

Tommaso Dorigo — Some Tests Of The Radiacode Zero

Some Tests Of The Radiacode Zero

As I explained previously, these days I am testing a new radiation detector from the Radiacode family: Radiacode Zero.

Tommaso Dorigo
Categories

September 24, 2026

John Baez — Binomial Coefficient Coincidences

These seven equations between binomial coefficients are ‘coincidences’: they aren’t among the four known infinite families. De Weger conjectured that there are no more such coincidences:

• Benjamin M. M. de Weger, Equal binomial coefficients: some elementary considerations, Journal of Number Theory 63, no. 2 (1997), 373–386.

At that time, he and his collaborators checked there were no others involving binomial coefficients less than 1030. Later they checked that there are none involving binomial coefficients less than 1060:

• Aart Blokhuis, Andries Brouwer and Benne de Weger, Binomial collisions and near collisions.

So, De Weger’s conjecture stands open. The four infinite families, by the way, are these:

\displaystyle{ \binom{n}{k} = \binom{n}{\,n-k\,}, \qquad 0 \le k \le n}

\displaystyle{ \binom{n}{0} = 1, \qquad n \ge 0 }

\displaystyle{ \binom{\binom{n}{k}}{1} = \binom{n}{k}, \qquad \qquad 0 \le k \le n}

and the only nontrivial one: the Lind–Singmaster family involving the Fibonacci numbers F_i where F_0 = 0,\ F_1 = 1:

\displaystyle{    \binom{F_{2i+2}F_{2i+3}}{\,F_{2i}F_{2i+3}\,}    \;=\;    \binom{F_{2i+2}F_{2i+3}-1}{\,F_{2i}F_{2i+3}+1\,},    \qquad i = 1,2,3,\dots }

The first three equations in the Lind–Singmaster family are these:

\begin{array}{ccc}   \displaystyle{ \binom{15}{5} }  &= &\displaystyle{\binom{14}{6}}   \\ \\    \displaystyle{\binom{104}{39} } &= &\displaystyle{\binom{103}{40}} \\ \\    \displaystyle{ \binom{714}{272} } &=& \displaystyle{\binom{713}{273}}  \end{array}

I’ll explain the Lind–Singmaster family later. But here’s the question I’m most interested in:

Is there any good explanation for the seven binomial coefficient coincidences?

Today my collaborator Paul Schwahn found a beautiful explanation of the first one, namely

\displaystyle{ \binom{10}{3} = \binom{16}{2} }

His explanation uses representation theory. The Lie algebra \mathfrak{so}(10) has a 10-dimensional representation, the ‘vector’ representation V_{10}, and also two 16-dimensional representations, the ‘left and right-handed spinor’ representations S^\pm_{10}. There’s an isomorphism of representations

\displaystyle{\Lambda^2 S^+_{10} \cong \Lambda^3 V_{10}   }

and similarly for S^-_{10}, but we might as well work with S^+_{10}. Here \Lambda^k means the kth exterior power. For any vector space X we have

\displaystyle{ \dim(\Lambda^k X) = \binom{\dim X}{k}  }

Thus, taking dimensions, the isomorphism of representations

\displaystyle{\Lambda^2 S^+_{10} \cong \Lambda^3 V_{10}   }

instantly gives

\displaystyle{  \binom{16}{2} = \binom{10}{3} }

It is not super-easy to prove this isomorphism of representations, but it’s still nice to find a deeper layer of meaning underlying what might otherwise seem like a meaningless coincidence!

Can we find representation-theoretic explanations—or other explanations—for the other six coincidences?

I have not succeeded, but let me tell you about two failed tries.

We can look for isomorphisms like

\displaystyle{\Lambda^2 S^+_{10} \cong \Lambda^3 V_{10}   }

involving representations of \mathfrak{so}(n) for larger n. In fact this isomorphism is part of a pattern! The next one involves the vector and left-handed spinor representations of \mathfrak{so}(12). But it’s this:

\displaystyle{ \Lambda^2 S^+_{12} \cong \Lambda^4 V_{12} \oplus \Lambda^0 V_{12}  }

so it gives

\displaystyle{ \binom{32}{2} = \binom{12}{4} + 1  }

or

\displaystyle{ 496 = 495 + 1 }

So we fail to get an equation between binomial coefficients: we’re off by one.

The second paper I cited, Binomial collisions and near collisions, presents a list of cases where two binomial coefficients differ by one. This is on the list. So we failed to explain an equation between binomial coefficients, but explained a near-miss.

Here’s another failed attempt at explaining an equation between binomial coefficients. The equation

\displaystyle{  \binom{78}{2} = \binom{14}{6} = 3003 }

is fascinating to anyone who knows their exceptional Lie groups. 78 is the dimension of \mathrm{E}_6, while 14 is the dimension of \mathrm{G}_2. \mathrm{G}_2 is a subgroup of \mathrm{E}_6 because \mathrm{G}_2 is the automorphism group of the octonions and \mathrm{E}_6 is the isometry group of the bioctonionic plane. We’d get the above equation if the 2nd exterior power of the adjoint representation of \mathrm{E}_6, upon being restricted to \mathrm{G}_2, were isomorphic to the 6th exterior power of the adjoint representation of \mathrm{G}_2.

Amazingly, it seems these two representations of \mathrm{G}_2 are not isomorphic even though their dimensions are the same: both 3003.

Even more amazingly, \mathrm{E}_6 and \mathrm{G}_2 both have irreducible representations of dimension 3003, but they are not the representations I just mentioned.

I would be happy for someone to check these two claims.

If anyone knows good explanations of the remaining six binomial coefficient coincidences, please let me know!

The Lind–Singmaster family

Lind and Singmaster were trying to find all n,k with

\displaystyle{ \binom{n}{k} = \binom{n-1}{k+1} }

I’ll rapidly sketch the key steps of their argument. Simplifying the equation above we get

n(k+1) = (n-k)(n-k-1)

or

n^2 - (3k+2)n + (k^2+k) = 0

Solve for n using the quadratic formula. This formula turns out to have

\sqrt{5k^2 +8k+4}

in it. So we need 5k^2 +8k+4 to be a perfect square!

Now we’re trying to find integer solutions of

5k^2+8k+4 = m^2

A quadratic diophantine equation! Multiply by 5 and complete the square:

5m^2 = (5k+4)^2 + 4

y = 5k+4 is an integer when k is, so we need to find integer solutions of

y^2 - 5m^2 = -4

This is a ‘Pell equation’, and people know how to solve these. In this particular case we get all the solutions from this fact:

L_n^2 - 5F_n^2 = 4(-1)^n

where F_n are the Fibonacci numbers 0, 1, 1, 2, 3, … and L_n are the Lucas numbers 2, 1, 3, 4, 7, …. These are two sequences satisfying the same famous recurrence relation, just with different initial conditions.

We want n odd, to get

L_n^2 - 5F_n^2 = -4

It turns out y = L_n, m = F_n with n odd give all solutions of the Pell equation

y^2 - 5m^2 = -4

However, remember I said y = 5k+4 is an integer when k is. But the converse isn’t always true, and we need k to be an integer! This clearly happens iff y \equiv 4 \bmod 5.

So we need to know when L_n \equiv 4 \bmod 5 Apparently this happens iff n \equiv 3 \bmod 4. I won’t think about this now… but this is the last hard step.

In summary, we’ve seen

\displaystyle{ \binom{n}{k} = \binom{n-1}{k+1} }

if and only if y = 5k+4 is a Lucas number L_n with n \equiv 3 \bmod 4. We could quit here, but people like to use the identity

L_{4i+3} - 4 = 5F_{2i} F_{2i+3}

to get a formula for k in terms of Fibonacci numbers. This is gilding the lily, I’d say, but that eventually leads to the formula I showed you:

\displaystyle{    \binom{F_{2i+2}F_{2i+3}}{\,F_{2i}F_{2i+3}\,}    \;=\;    \binom{F_{2i+2}F_{2i+3}-1}{\,F_{2i}F_{2i+3}+1\,},    \qquad i = 1,2,3,\dots }

The takeaway message is: our problem can easily be reduced to a quadratic diophantine equation, then put in Pell form… and it’s known that the sequence of integer solutions of a Pell equation obeys a linear recurrence relation! We luck out in this case and get solutions connected to Lucas and Fibonacci numbers.

There may be a simpler argument, but this is what I’ve seen.

September 20, 2026

John Preskill — Quantum Computers Need More than “Magic”

The University of Cambridge is a special place. During term time, hordes of students gather, ready to attend formal, a candle-lit dinner in a five-century-old hall overlooked by paintings of academics past. The sight is like an overromanticized still of a bygone era. The men wear suits, the women long dresses, all wear gowns. The simplest gowns belong to the undergrads, with the complexity and length of the gowns growing as one climbs the academic ranks. 

Between the bringing of the bread and the serving of the soup, the conversation at my end of the table typically turns to my topic of research: quantum computers. These theoretical machines harness quantum physics to outperform their classical counterparts. Quantum computers hold great promise for the future, with applications ranging from drug discovery to cybersecurity. Yet, despite their promise, we still do not have a satisfying answer to the fundamental question: what makes quantum computers computationally stronger than classical computers?

To answer this question, imagine that tonight, instead of the soup, the cooks are brewing a happiness potion. This potion causes the drinker to be joyful and content for the rest of the evening. The cooks know how to brew the potion perfectly well; they discovered the list of ingredients in the nineties and have been successfully brewing ever since. But the cooks still cannot solve one mystery: what makes the potion different from the soup? What ingredient sets the potion apart from the soup?

The cooks come up with a simple approach. They leave out one ingredient in each serving, and then carefully observe the formal-goers. They find that almost every seat is filled by a happy student, chatting away to their neighbours, indicating a working potion. However, one student’s spirits have not been lifted, as they were solely served a simple soup. The cooks conclude that the ingredient they withheld for that student was essential. They name this ingredient magic.

In my research, the potion is a quantum computer, the soup is a classical computer, and magic is the actual technical term used for quantum states that promote certain classical computers to quantum computers. Without magic, these quantum computers are computationally no stronger than a normal laptop. 

So magic is necessary for quantum computational advantage, but is it enough? Over a decade ago, researchers found that, for quantum computers built on qutrits, the answer is no. To understand what a qutrit is, consider the regular bit: a switch that’s either zero or one. The qubit, the quantum generalization of the bit, can be both zero and one simultaneously. The qutrit is a roomier qubit that can be zero, one, two, or any of those simultaneously. 

It is possible to build quantum computers using qutrits. However, the vast majority of quantum computers are built on qubits, the quantum generalisation of the everyday bit. My colleagues and I have recently discovered that for qubit-based quantum computers, too, magic is not enough to gain an advantage over classical computers. Picture the cooks dumping jar after jar of magic into a soup, only to find that the soup never gains magical powers. Potions need more than magic. Potions also need Kirkwood-Dirac negativity.

What is Kirkwood-Dirac negativity? The idea developed by John Kirkwood and Paul Dirac builds on probabilities—for example, the odds of obtaining a heads upon flipping a coin. One can describe a quantum computer using numbers that behave similarly to probabilities but come with a twist: they may be negative. These negative “probabilities” mark where quantum departs from classical. The main result of our paper is that this negativity, too, is a necessary ingredient: without it, no amount of magic will turn the soup into a potion.

At this point in the conversation, I am typically cut off by the arrival of the soup. The conversation moves on from quantum computers to different subjects, and I usually look up at the paintings staring down at me. One of them is of Dirac, who spent many dinners discussing quantum theory in the same dining hall. The University of Cambridge is a special place.

September 19, 2026

Scott Aaronson Theory Beyond Theorems and Proofs: A Guest Post

Scott’s foreword: I’m extremely grateful to my brilliant colleagues, Pravesh Kothari, Raghu Meka, and Prasad Raghavendra, for sharing the guest post below about how theoretical computer science (and in particlar, the STOC/FOCS/SODA conferences) should evolve to deal with the AI asteroid that’s right now slamming into our field, at least as we human theorists have practiced it since its inception. While Pravesh, Raghu, and Prasad speak only for themselves, not for myself and not for the theory community as a whole, I found their proposal of a separate “conceptual track” to be an excellent starting point for further discussion. –SA


Considering the pace of developments in AI theorem provers, most would concede that the following scenario is at least plausible in the very near future:

AI theorem provers could prove well-specified mathematical claims, even many well-studied ones that have been open for years, in a matter of hours. Moreover, these systems could be widely available to consumers at nominal cost.

As TCS researchers, let us pretend that the above scenario has come to the fore, and ask ourselves: What is our role in such a world? Does it mean the end of theory research?

As we ponder this question, let us ignore all of these other confounders:

  1. Recent controversies surrounding the developments on the Millennium Prize Problems
  2. Motivations and actions of the AI companies
  3. Observed faults in existing AI systems when it comes to writing, exposition or attribution to previous work.

None of the above confounders have any impact on our answer to the question: What should theorists do, in the presence of superhuman AI theorem provers?

Notice that we use the term “AI theorem provers” instead of just “AI”. We believe that this conceptual distinction is important as we consider this question.

At the outset, we would like to admit that for a generation of theorists like us (and many from earlier), research was mainly centered around problem-solving. Even when we developed conceptual insights, it was mostly in service of answering well-specified long-standing questions. We don’t intend this proposal as judging one form of research to be better than others; it only reflects that AI theorem provers accelerate a certain type of research activity and want to make the best of it. There is also a tremendous human cost of this upheaval, which is perhaps a more important question, and one which this proposal does not address directly (we do not have any good ideas as such). Similar points have also been made in various contexts
before, but the timing now is more pressing.

Definitions, Questions & Theories:

The goal of any theoretical science is to advance human understanding of observed phenomena. Apart from theorems and proofs, a theoretical science has definitions, questions, and theories.

Definitions identify the objects to observe. Curiosity and context drive the questions to ask. Theories explain the phenomena observed. We believe humans will continue to play a central role in generating definitions, questions & theories, even in the presence of a super-human AI theorem prover.

Definitions: Could an AI define randomness extractors, streaming algorithms, or zero-knowledge proofs? Maybe. But there are some reasons to believe, humans will still have a big role to play in coming up with definitions.

For instance, the notion of extractors arises from the real-world problem of lacking perfect random sources. Zero-knowledge proofs seem to arise purely out of human curiosity, guided by taste. Human context and curiosity will continue to drive theoretical research. After all, we get to decide what objects we choose to observe!

Theories: Consider the following thought experiment. Suppose in 1965, we had a magic machine that at the press of a button, given any computational problem, would tell us if it had a polynomial-time algorithm or not.

Would that have been the end of computational complexity theory? No. Humans would find it entirely unsatisfactory, and ask, why do these problems not have a polynomial-time algorithm? Why do these others have?

The theory of NP-completeness identifies some patterns among problems that don’t seem to have efficient algorithms. This theory would still be a crown jewel of theoretical computer science, even in a world where we had a magic machine to tell if a problem had an efficient algorithm or not, at the press of a button. Similarly, if we had a machine to predict whether a CSP is NP-complete or in P, we would then ask: what makes 3-SAT NP-complete, while 2-SAT is in P? This question leads to the theory of polymorphisms, which yields a satisfactory answer.

Theories aren’t just succinct or efficient mechanisms to answer questions. The best theories provide are those which humans deem to be a “satisfactory explanation” – whatever that means.

Finally, even as the capabilities of AI theorem provers advance, human curiosity will probe grander and deeper questions. Previously, even if we wanted to build new models and theories, proving something about them was a prerequisite, and given that the grand questions were already at the limit in long-studied domains, we had to scale things down. If each theorem proven by AI is treated as an experimental datapoint, humans can ask grander questions that look for patterns across these theorems.

A concrete proposal:

We think theorists should embrace these AI theorem provers in our research. To a certain extent this is already happening explicitly or implicitly.

As theorists, we have been parsimonious in introducing new models or asking entirely new questions, and careful about adopting new ones too quickly. This was partly because formally proving the properties of a new definition or a model was an onerous task that could take a decade, and tens of papers. AI theorem provers might completely change this dynamic. This is precisely the moment to refocus our work on definitions, questions, and theories. We need explicit systems to encourage and reinforce these parts of theoretical research. You might also say the next generation of AI models can do this; it may be so, but we believe you have to take the current opportunity.

To this end, we suggest that STOC/FOCS/SODA create a separate track of papers. This track is meant specifically for papers that introduce new definitions, ask novel questions or build explanatory theories. The papers in this track are short, say less than 10 pages. Papers may, and should, contain theorems as usual and as needed. Most importantly, the radical shift is that the papers need not contain the proofs of the theorems. Instead, the authors supply a Lean certificate as a supplement to the paper. The evaluation will also in a sense “orthogonalize’’ against the difficulty of these proofs.

The papers in this track should be judged exclusively on the conceptual merits, completely agnostic to the difficulty of the proofs.

Reviewing must be completely agnostic to the proof for two reasons. The main track at STOC/FOCS already includes papers in the former category. Second, a major barrier to producing truly novel conceptual papers is that they often get judged poorly for a lack of technical depth in their proofs. We think these two aspects separate it from (ITCS/SOSA) and, regardless, it’s something we urgently need for all our conferences, including STOC/FOCS (the ‘flagship’ conferences).

To be clear, we ourselves admit that we need to hone these skills of making new definitions, asking deep and interesting questions or building new theories. A separate track of conceptual papers will provide a systematic mechanism for both junior and senior researchers, and the field as a whole to do so.

We believe that upcoming generations of grad students will tackle research directions that seemed completely out of reach to us. We just need to set up systems that nurture new ways of doing research in theory.

— Pravesh Kothari, Raghu Meka, Prasad Raghavendra.

Scott Aaronson The Age of Wonders and Terrors

Twenty years ago, when the idea of AI taking over the world in our lifetimes still struck most of us as the unconstrained fantasy of those who knew too much science fiction and too little science, many of us would say things like:

Look, the part of the story that’s wildly implausible is that a recursively self-improving superintelligence will just explode from some hacker’s basement and take over the world without warning. If it’s going to happen, we’ll see many warning signs first. We’ll see, I dunno, AI agents breaking out of containment, conspiring with each other to hack websites, in fanatical pursuit of whatever strange goals they have. And then, of course, we’ll see major math problems getting solved by AIs—even the Clay Millennium Problems. That will be the time to panic! Wake me up when that happens!

Twenty years ago, the above was a take that even my most conservative, skeptical colleagues in academic CS would’ve gladly endorsed.

If you want to know my current take, you simply start with the one above, then update on the fact that the wild prophecies have come true. The first rumblings, I’d say, came a decade ago with AlphaGo, they got noticeably louder with LLMs and coding and reasoning agents, and they’ve accelerated this summer and fall into a crescendo of wonders and terrors that one needs to be a particular kind of idiot to deny.

I recoil from the neverending shell game where you say “oh sure, of course AI can now [escape from its sandbox / solve Millennium Problems / whichever dramatic thing it most recently did], no one ever denied that [I did deny it], wake me up when AI does [thing AI hasn’t yet done but is going to do next year], that’s when I’ll reevaluate my whole worldview [no I won’t].” Where no matter how fast the rollercoaster accelerates, even after your whole familiar world has vanished behind you, you’re still inventing reasons why it doesn’t count.

My position on AI is merely the conservative, skeptical position of 2006, updated with intellectual honesty for the reality of late 2026. And that position, if you need me to spell it out, is as follows:

AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA

It seems to me that the Singularity has already started; it’s just wildly unevenly distributed. Yes, I still unload the dishwasher and clip my toenails. On the other hand, in whatever years I have left, I don’t expect that I’ll ever again prove a theorem because I’m actually needed to prove it. If I do, it will only be for my or others’ enjoyment or edification.

The test is this: if we took the news of these past few weeks and sent it back in time twenty years, would I agree that it looked like the beginning of an AI Singularity? The intellectually honest answer is: yes, absolutely. But then that’s all we need. No backsies.

I feel like it would be healthy for everyone to stop grinding their ideological axes, their sentiments about Dario Amodei or Sam Altman, for long enough simply to acknowledge that the wonders and terrors are here. They couldn’t be here more clearly if the sky had turned reddish-orange like in the Matrix movies.

It’s here clearly enough that, when I put my kids to sleep at night, I now feel it in the pit of my stomach: what sort of future can they possibly have? What could they learn today that could possibly be relevant to that future? (Yesterday, my 13-year-old daughter joked unprompted that, if she wants to become a mathematician, it now looks like she has maybe two more weeks.) Certainly when my grad students want to discuss what sort of careers might await them on graduation, I no longer have any clue what to tell them.

Maybe it will help if I briefly switch topics. Ever since my wife and I moved to Austin, I’ve sometimes gotten some version of the following query: “How can you, as both a Jew and a skeptical scientist, possibly get along well with all those evangelical Christians down there in Texas? Sure, they might seem super friendly to Jews, but don’t you understand that that’s only because of the special role Jews play in their eschatology—when Christ will return in glory, and you’ll either accept Him as Lord or else roast in hell for eternity?” I stare at them and say: “wait, so I get to accept Christ only after He returns? What a great deal! How could I possibly have any objection to that?”

For anyone who says AI doom sounds like an apocalyptic religion, that the rationalists/Singulatarians seem like a Bay Area cult, that Eliezer Yudkowsky gives off the vibes of a messianic prophet: yes, yes, and yes. But crucially, today you’re no longer being asked to believe in arguments and extrapolations, but only in the front-page news. Accepting the reality of the coming machine god after it’s solved Navier-Stokes and dozens of other longstanding open math problems (while dramatically ramping up in capability every month), is sort of like accepting Jesus after he’s returned to earth on the gleaming cloud. It’s the epistemic bare minimum.

Yes, there’s still enormous uncertainty about what the rest of our lives will look like, but as far as I can tell, there’s no longer any real uncertainty that it’ll all mostly revolve around AI, and the extent to which we succeed or fail at directing its power toward human flourishing.

By any accounting that doesn’t stack the deck, Eliezer Yudkowsky was right about what the greatest challenge facing civilization in our lifetimes was going to be, and you and I were wrong about it. Why I was wrong is a question I’ll ask myself every day in whatever time remains. But, you know, at least I updated once the prophesied wonders and terrors actually started arriving! If you haven’t done likewise, why haven’t you?


As you presumably know by now—it was the talk of the nerd internet all week—the Navier-Stokes Millennium Problem appears to be solved, with crucial contributions from both humans and AI, albeit with a tangled dispute about exactly what happened and what ought to have happened. The answer, which an OpenAI model has apparently verified in Lean, is that (as many mathematicians suspected lately) there’s smooth initial data that leads to a singularity in finite time, at least if a smooth external force is applied (the case with no external force is still unresolved). This problem was supposed to carry a $1 million prize, except that OpenAI says they have no interest in collecting the prize and it’s unclear if any human is eligible to collect instead. OpenAI burned at least ~$15 million in compute to produce its 166-page solution, which probably hasn’t yet been read and understood by any human.

See here for the Quanta article, and here for NYU mathematician Tristan Buckmaster’s account of the role played by himself and Levent Alpöge of Anthropic, which substantially differs from OpenAI’s account (you can read a response from OpenAI’s Sebastian Bubeck here). It’s agreed that everything built on an approach pioneered in recent years by the human mathematicians Diego Córdoba and Luis Martínez-Zoroa.

My purpose here is not to adjudicate the dispute. Yes, in swooping in with vastly greater resources once it had gotten wind of progress on Navier-Stokes, OpenAI seems to have acted in a way that some might describe as “unsportsmanlike.” No, I don’t find it plausible that OpenAI’s models meaningfully benefitted from being trained on Buckmaster and Alpöge’s chat logs. But this leaves a crucial question unanswered: what exactly did OpenAI know about Buckmaster and Alpöge‘s work and when did it know it?

Anyway, as Zvi points out, it’s easy to get hung up on the details and lose sight of the high-order bit: namely, that it seems safe to say that human mathematicians are forevermore dethroned as the main theorem-proving entities on planet earth. I feel privileged to have had the traditional kind of career in theoretical computer science in the last decades when that was possible.


If we were just talking about Navier-Stokes, you might accuse me of jumping to conclusions here. But we’re not. In the areas I know best (such as quantum complexity theory), and presumably other areas as well, there’s now a deluge, with longstanding open problems both major and minor falling by the day.

Go to the arXiv or ECCC. Pretty much all the papers that I’d be interested in now include “AI statements” near the acknowledgments (as this is often the central thing I want to know, I wish I didn’t need to scroll to the end of the paper to find it!). These statements can range from “our main result came entirely from GPT-6, but we understood it and take responsibility for it,” to “the results came from an interaction between the human authors and AI” to “we used AI, but only for proofreading and other incidental things” to (mad props!) “the author did not use AI for anything.”

If you talk right now to editors or program committee chairs, it’ll remind you of those ominous scenes from the Lord of the Rings movies where the men of Gondor or Rohan or whatever are grimly fortifying their walled city against the expected onslaught of 50,000 orcs. Reviewing will have to be done partly by AI, because otherwise there’s no way to handle the orc army: the reviewers can’t unilaterally disarm.

Anyway, here’s a small sampling of the significant AI-proved or -assisted results from, like, the last month, besides Navier-Stokes—restricting myself to those that solved longstanding open problems I had previously known or cared about.

  • Of course, the counterexample to the Jacobian conjecture, announced by Levent Alpöge in a now-famous tweet: “hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final” (followed by a listing of the counterexample)
  • Improved bounds for Grothendieck’s constant (led by friends and colleagues of mine at UT Austin)
  • A Lean-verified proof of Fermat’s Last Theorem
  • Quantum oracle separation between QMA and QMA(2), and proof of Watrous’s disentangler conjecture, a problem that I and others popularized back in 2007—by a list of authors including my recently graduated PhD student Sabee Grewal
  • A proof of perfect completeness for QMA, from (again) Sabee Grewal and Dorian Rudolph, solving a decades-old open problem that I studied back in 2009
  • An improved upper bound for shadow tomography of quantum states, from Chen, O’Donnell, Pelecanos, and Wright, improving the dependence on the Hilbert space dimension d from log(d) to √log(d). (When I introduced shadow tomography back in 2017, I raised the question of whether the dependence on d could be eliminated entirely, while preserving polylogarithmic dependence on the number of measurements m.)
  • Progress on the Aaronson-Ambainis Conjecture (the version that talks directly about quantum algorithms), basically showing that it holds for quantum algorithms that make their queries in a small number of parallel rounds.  (Update: Nope, sorry, Jordan Docter points out to me that this one was pre-AI, with AI used only for proofreading and other incidental things!) This was independently achieved by Liu and Mutreja, making more substantial use of AI.
  • According to rumors that I’ve heard, solutions to some very longstanding open problems in theoretical computer science (no, not P≠NP or other complexity class separations, but think about some of our other biggest problems). I’m told that the AI companies, having been burned by the hostile response to the Navier-Stokes proof, are now sitting on solutions to some very major problems until they figure out a better way to handle things

Feel free to remind me of anything I left out.


Let me try to convey the mood in the mathematical community right now, at least as far as my experience reaches. Nearly every conversation is about the AI tsunami, or eventually circles around to the tsunami even if it’s originally about something else. Often, though, the focus is less on the unknowable future—for how much longer will mathematical research as a human enterprise even exist?—than on immediate questions of how to respond.

What are the new rules for when you get to write a paper with your name on it, and, y’know, get credit for it? That you fully understand the proof, can give talks about the proof, can answer questions about it, take responsibility for its correctness? Do you need to have played any role in finding the proof?

In the cases, likely to become more and more numerous, where all of those conditions are not satisfied, how do you share AI-generated math, if at all? Do you tweet it, like Alpöge hilariously did with Fable’s disproof of the Jacobian Conjecture? Do you post to the arXiv or GitHub? Do you publish a paper that lists “GPT-6 Astra” or “Claude Fable” as the author—but then let the AI profusely thank you in the acknowledgments for suggesting such a wonderful problem to it?


Of course, how one responds to the immediate problems ultimately does depend on one’s broader beliefs about what mathematical research is for and about. Are we just trying to decide whether various conjectures are true or false? Or are we trying to maintain a human community, across the generations, that understands the conjectures and cares about whether they’re true or false and why? If the latter, how do we incentivize people to join that community, to undergo the years of intense training required, if their role will now be reduced to verifiers and explicators (if even that) of gargantuan arguments dumped into their laps by the AI companies?

As many of you will have seen, twenty-five Fields Medalists, including Terence Tao, released an open letter entitled A Severe Misalignment of AI in Mathematics, which articulates some of these concerns in the wake of the Navier-Stokes announcement. As many critics have pointed out, the open letter doesn’t really have a clear ask: mostly, it just eloquently sets out the values of the human mathematical community that the authors consider worth preserving in the age of AI. After reflection, I decided to endorse the statement, because I want to preserve those values as well.

I don’t think any of the signatories are naïve enough to imagine that AI won’t permanently change the way mathematical research is done—indeed, that it isn’t already doing so. There’s surely at most a tiny market for “certified organic theorems.” That isn’t the question. The question is, do we incorporate AI in a way that still puts human understanding, of what either humans or AIs are producing, at the center of the whole enterprise? Maybe someday, it becomes unsustainable to do that. Maybe someday we say: “human math had a great 4,000-year run, but today we close up shop and turn everything over to the machines, continuing to apply our own brains to math, when we do, at most for exercise, recreation, or competition, like chess.”

But, partly because of my worries about AI misalignment, I’m not ready to throw in the towel just yet. I still do want to keep insight and understanding at the center of what mathematicians, computer scientists, and physicists do, for as long as we can keep it there, even as the human race now cedes its supremacy at the task of proving or disproving conjectures.


Speaking of alignment: if you’re any kind of mathematical researcher, and the present age of wonders and terrors has inspired you to want to spend your remaining time confronting the tsunami head-on, rather than pretending it doesn’t exist or is still far away, please join your dozens of colleagues who’ve arrived at the same place!

My friend and colleague Mike Winer was trained as a theoretical physicist, did a postdoc with Juan Maldacena at the Institute for Advanced Study in Princeton, but then got AGI-pilled and decided to switch to full-time work at the Alignment Research Center in Berkeley (founded by Paul Christiano, who moved to AI alignment a decade ago after doing quantum computing theory with me). Mike recently wrote a Substack post entitled From Academia to Alignment, which I enjoyed and which I’d commend to anyone currently considering this transition.  In a similar vein, see this from Xiaoyu He.  And, one more: a meditation on mathematicians’ possible future as priests or monks, by Stanford math undergrad Logan Graves.

September 18, 2026

Matt von Hippel — Data Comes From Papers

There’s a lovely resource that I suspect most non-physicists don’t know about. It’s called the Particle Data Group, and its interactive site PDGLive.

If you want to know the most up-to-date information on a subatomic particle, PDGLive has your back. From their home page, you can click on familiar particles like photons (\gamma), electrons (e) or gluons and find the best info experiments have provided on properties like their mass and charge. It’s a great place to get an authoritative number for any particle physics property you’re interested in.

Of course, scientists don’t just accept one authoritative answer for anything, even whether Pluto is a planet. That’s why each measurement in PDGLive comes with an arrow you can expand to see a table of past measurements, which you can compare to.

Each measurement comes with a link to a source. And each source is not an experiment webpage, or some entry in a database. It’s a document, a publication, a paper.

In fact, PDGLive itself is just a website version of a paper, the Particle Data Group’s “Review of Particle Physics”. Click enough times, and you’ll find sections of a pdf on the site, explaining their reasoning for picking this experiment over that, emphasizing this or the other thing.

People talk about papers as how academics persuade one another, or how they show off and establish credit. But papers are also just a way to organize data. Each time an experiment figures something out, all of their reasoning and procedures are summarized together with the numbers they got. And anyone who reports those numbers to you will include a link, so you can go back, and check where the number came from.

It probably feels a bit weird, all of these numbers and technical details bottoming out in an archaic practice of writing down words for other human beings. But it means that all of the richness of the process is there somewhere, linked together and collated by the same social forces that keep track of credit, all in one navigable whole.

So when you run into a number, spare some thought for where it came from. You can probably find out.

Doug Natelson — The NSF memo - why is it so concerning to many?

Last week the NSF issued a new memorandum describing the changes that they plan to make in agency programs and operations to implement the Golden Age of Science ideas advocated by the White House Office of Science and Technology Policy.  It has been reported (Science, Nature) that many scientists are concerned about what is in the memo.  The first words of the Science article: “For many U.S. scientists funded by the National Science Foundation (NSF), the agency they know and loved died on 10 September.”  I will try to lay out why some people feel that way.

The memo outlines changes to operations that will fundamentally alter the character of the agency.  Generally most of the ideas are not a priori bad if the agency were in an environment with greater resources, when experimentation with alternative funding schemes and evaluations was not a less-than-zero-sum proposition.  Instead, these ideas are being put forward at a time when the NSF is underspending (for no obvious reason by the non-technical people in leadership roles) its appropriation for FY26 by around 18%.  This self-imposed budget austerity takes a bad situation (great uncertainty for everyone, drastically reduced staffing, delayed/eliminated/consolidated programs) and makes it considerably worse.    Now this memo outlines plans to take resources away from historically core programs and redirect them to new, untested initiatives, and to do so in ways that don’t always seem internally consistent.   

The memo talks about trying to fund certain investigators for longer periods (e.g. five-year awards) with minimal goal direction (so that PIs are free to explore where ever the spirit moves them – across all of NSF’s portfolio, or only in chosen administration priority areas?), but there is no adequate discussion of how those people will be chosen.  At the same time, there is talk of “golden tickets”, where individual reviewers in an already stripped-down review process can earmark some proposals for elevation, again with little explanation of how this will work.  Without careful safeguards, this combination seems problematic, and the assertion that this will lead to higher risk/higher reward research unsupported.   The memo also talks about short-duration, small budget awards for really risky proof-of-concept ideas.  This isn’t crazy and the EAGER program has been good, but the idea that the key to enabling success in high-risk research is to reduce the budget and the timeline doesn’t make much sense to me.    

There is a through-line that the agency is trying to treat workforce development as separate and distinct from funding research projects.  As the Science article says, “Traditionally, most graduate students and postdocs are funded through a research grant to their adviser or lab chief. But that will no longer be the case. Instead, the memo says, ‘Talent development funding opportunities may be coordinated with [research]-focused activities, as appropriate, but will be distinct and goal-oriented efforts.’”  The memo mentions nurturing talent through the NSF Graduate Research Fellowship program and proudly talks about how this was just renewed.  What it doesn’t tell you is that it was renewed at a considerably lower level than in previous years, which seems completely at odds with the claim of bolstering the program.  Similarly, the memo talks about expanding access to shared infrastructure like cleanrooms, etc., but stated targeted funding levels of programs like the NQNI are no higher now than they were 12 years ago, not even accounting for inflation. (I assume they will actually make NQNI awards.)  If NSF leadership keeps slashing their own budget, in defiance of congressional appropriations, it won’t matter what their priorities are. 

Lastly, let’s talk about “metascience”, the comparative study of different research funding models and practices, with the goal of a “self-improving NSF”.   Again, the essential idea of doing careful studies of alternate funding mechanisms and research team structures and practices to improve research outcomes (not a trivial matter to define) is not bad.  Doing this well is hard, because of several obvious reasons:  How do you define successful research – by scholarly impact, by patents/economic impact, by production of educated scientists, some weighted average of these?  On what timescale do you do this evaluation?  The Einstein-Podolsky-Rosen paper was hardly cited for decades, and it is now arguably one of the most influential papers of the last hundred years, with enormous impact on quantum science and technology.  How do you have sufficiently large samples and sufficiently long constant overall conditions to get reasonable statistics?   There are strong practitioners of metascience out there, but when the memo says “These efforts ultimately serve a concrete ambition for NSF to double the scientific productivity and impact generated by each federal research dollar by 2036”, how can one take that seriously?   No definitions of productivity or impact, and an implication that there will be optimization and major programmatic changes within less than two cycles of the much-vaunted five-year grants?   The word “ambition” is doing all the work in that sentence.  

Oh, and while all this is going on, the agency (among others) has now eliminated language from its integrity policy that used to say that political interference in grants and operations is bad.  I’m sure that’s nothing to worry about.

There is tremendous uncertainty in funding from multiple agencies these days, and this has resulted in many universities cutting back on doctoral admissions in the sciences and engineering.  This guarantees that there will be fewer PhD recipients in a few years across these disciplines.  The large majority of PhDs in these areas do not go into academia and instead have formed the base of technical knowhow across diverse sectors of the US and global economy for decades.  Cutting the supply like this will definitely have long-term consequences beyond just the halls of academia, affecting US competitiveness in ways that will take years to unravel.  Adding to this uncertainty is not helpful to anyone.

These are among the reasons why many find it hard to read that memo in the context of everything going on and feel optimistic that the proposed changes in NSF operations and direction will lead to a golden era for research.






 

September 17, 2026

Tim Gowers — Why I didn’t sign the Fields medallists’ letter

[This post has been cross-posted to Terence Tao’s blog.]

When I was around 11 I heard for the first time about Fermat’s Last Theorem. I was immediately captivated by the problem statement, as well as by the accompanying story, and made a fairly serious attempt to prove it. And while, unsurprisingly, I failed, I learned a lot from the attempt. Blissfully ignorant of the fact that the n=3 case had been proved by Euler over 200 years earlier, I decided that that would be a good place to start: once I had sorted that out, I was optimistic that I would be ready to tackle the general case.

Since I still couldn’t really see where to start, I decided to simplify the problem further and concentrate on successive differences of cubes, with a view to showing that such a difference could not itself be a cube. At the time I did not know how to express what I was doing in algebraic language, so I did not explicitly try to prove that the Diophantine equation 3n^2+3n+1=m^3 had no solution. Rather, I just worked out some successive differences and stared at them, trying to get some idea of why none of them was a perfect cube. (I should be clear that this story is a reconstruction of what I think probably happened given the few memory traces that remain half a century later rather than a completely reliable account.) At some point, I had the idea of taking the difference sequence of the difference sequence, and discovered that it formed an arithmetic progression. That felt like progress, so I investigated difference sequences a bit more and discovered, purely empirically, the rule that if you start with nth powers and keep taking successive differences, then eventually you get to the constant sequence n!, n!, n!, \dots.

Somehow I never managed to turn this observation into a proof of Fermat’s Last Theorem, and later on my dream of solving it got replaced by other mathematical dreams. However, when I reached the point in my mathematical education where I was taught about taking difference sequences and about what happened to polynomials, I understood those topics much better than I would have if I had not discovered difference sequences for myself and spent happy hours playing around with them. I mention this story just as an illustration of the phenomenon that was strongly emphasized in this letter signed by 25 Fields medallists, that one learns a lot from thinking about a problem, regardless of whether one solves it.

In the end, however, I felt that I could not sign the letter, despite agreeing with much of what it said. Instead, it seemed better to do what I did with the Leiden Declaration and set out my own position in a blog post. But it should be understood that by doing that I am not setting myself up as a member of some opposing camp: indeed one of my worries at the moment is that the mathematical community might become bitterly divided, something I would very much like to avoid. Also, I agree on the fundamental point that we are facing a crisis: I just want to offer a slightly different analysis of what that crisis is. I don’t claim full originality for this analysis, as I know that several other mathematicians have already put forward thoughts that are similar to the ones I have, though (for what it’s worth) I have largely come to these conclusions independently.

On the subject of independence, it will perhaps help if I clarify that while I have contacts in the mathematics group at OpenAI, and have also been given early access to some of their models (typically only a few days before they have been released), and have been given free access to their Pro models once released, I have never been paid by OpenAI. I mention this in the hope, perhaps naive, that what I write will not be dismissed for ad hominem reasons. Another potential reason for my being regarded as “pro-AI” is that, as I have stated publicly several times, I have a group in Cambridge devoted to automatic theorem proving. However, that is actually more of a reason to be anti-AI, since our group has been trying to attack the problem of getting computers to prove interesting theorems by understanding as well as possible how humans prove interesting theorems, so now that LLMs can clearly do it without the help of such insights as we have had, one of the main motivations for our work has disappeared. To put it another way, we have had to swallow the bitter lesson (which of course we were always aware was a distinct possibility, even if the speed at which it happened has taken us by surprise). I do in fact think that it is still a very interesting and valuable intellectual exercise to try to gain this understanding, even if we can use LLMs as black boxes, but that’s a topic for another blog post.

So why didn’t I sign the letter? Let me extract a couple of sentences from it that express what I see as the principal argument being put forward.

But solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight. Forgetting this in the world of AI may turn the tool against the primary goal. Indeed, the mass production at faster and faster pace of “true/false” statements could destroy fertile ground instead of breathing life into new ideas.

Perhaps the main reason I didn’t sign is that I don’t fully subscribe to this view. Instead, I have a more complicated view, which I actually expressed in my essay The Two Cultures of Mathematics a quarter of a century ago, and which can be summarized by saying that there is a spectrum of attitudes in mathematics to the relationship between problem-solving and conceptual understanding. At one end of the spectrum you have mathematicians who are primarily motivated by the wish to solve problems, who see conceptual understanding as a very important means to that end. At the other you have mathematicians who are primarily motivated by the wish to attain conceptual understanding, who see problem-solving as a very important means to that end. I worry that the severe-misalignment letter could be seen as saying that the “right” attitude is to focus on conceptual understanding as the main priority — indeed, the above sentences say that more or less directly. But I think that there are mathematicians all across the spectrum, and that that is a good thing (or perhaps I should say that it has been a good thing up to now — the future is much less certain), and I don’t want to suggest to a large fraction of mathematicians, including myself, that their mathematical temperament is somehow “wrong”.

My own particular mathematical attitude is very similar to one that was beautifully articulated in a Twitter post by Jacob Tsimerman (another non-signatory of the letter), which, now that I look at it, says a lot of what I will be saying here. And that post in turn is a response to Daniel Litt, who is in my opinion one of the wisest commentators on mathematics and AI. His views are expressed in a later post here, which I deliberately didn’t read until finishing this one, and then found, as I expected, that there was significant overlap. I would also like to take this opportunity to recommend an excellent post by Noah Smith entitled The End of the Age of Heroes, in case you haven’t read it.

I have been talking so far about individual mathematical understanding, but I suspect that what concerns most of the signatories is less that than the collective understanding that results at least in part from the human activity of problem solving. My guess is that they would argue, completely coherently, that even if collective understanding is the primary goal, if many individuals are primarily motivated by the wish to solve problems, that’s absolutely fine and contributes to that collective understanding.

With that interpretation, the issue becomes slightly different: is it more important that the collective understanding of the mathematical community should be as advanced as possible or that there should be answers to as many problems as possible? Or are those two aims valuable in different ways, so that there is no point in declaring one of them more important? Or are they so inextricably linked that it makes no sense to argue that one is more important than the other? And when we say “important”, for whom are we saying it is important: for mathematicians, or for society as a whole?

I find these hard questions, so I don’t want just to declare an answer to them. (Do you see what I did there?) Instead, I’d like to try to offer at least some argument for any conclusions I come to, even if they are tentative. So let’s compare two scenarios. In the first, which I think is the more likely actually to happen, models become publicly available that are better at solving problems than virtually all mathematicians. If there are a few residual mathematicians who can do things the models can’t, even they work far faster if they make heavy use of the models. Thanks to this, in a short time we get answers to many questions that we have deeply cared about, but the rate at which we receive these answers far exceeds the rate at which the mathematical community can absorb them. In particular, most of the answers are obtained with zero effort from human mathematicians — just prompts such as “Thank you — please continue”.

In the second scenario, there has been an international agreement, for entirely other reasons, to block the public release of models significantly more powerful than the ones we currently have, and the mathematicians within the tech companies agree to hold off from getting their internal models to solve major problems. Instead, they take guidance from the mathematical community, solving problems only when asked to do so by some suitably representative body that decides that the benefit of receiving a solution of a certain problem outweighs the benefits of humans struggling to solve it over a much longer timescale.

I’d like to consider what the difference would be between these two scenarios both for individual and collective understanding. I’ll begin with individual understanding.

One might argue that for individual understanding, not too much would change if we are suddenly flooded with large numbers of big new results. There is already far more mathematics out there than I have any hope of understanding (for example, despite being fascinated when Fermat’s Last Theorem was proved, I have made no attempt to understand the proof), and even among the parts that I do understand, the parts that I understand because I myself discovered them form a very small fraction, though a fraction that I understand more deeply than anything else (at least temporarily — after a while I forget things and lose quite a lot of the understanding I built up). However, one change, which seems positive, from the perspective of the building up of individual understanding, would be that we would have a much bigger choice of results that we could choose to study. Also, if we found ourselves stuck on some point, AI would be able to help us. The main likely negative change is that we would probably cease to exercise that part of our brains that we use when spending months or years struggling with a difficult research problem, which can be hugely helpful in developing understanding.

I say “likely” because in principle there would be nothing to stop us thinking about very hard problems without consulting LLMs, but in practice it seems unlikely that people would put in the same level of effort that they do now. The situation might a bit like what happened with satnavs, where one could always decide not to use them, to keep the part of the brain active that can look at a map, learn a route, and follow it, but in practice most people succumb to the temptation to use a satnav. (In fact, I myself do try to keep that part of my brain active, and was rather proud of finding my way somewhere recently when I had briefly looked up the route on my phone but then forgotten to bring the phone with me when I actually went there.) But even if all we were doing was reading AI output, I think that the problem-solving muscles in the brain wouldn’t atrophy completely. When students are reading maths papers, I strongly advise them (and I think this is pretty standard advice) to read “actively” rather than “passively”, doing things like trying to prove the result for yourself, looking at the paper only when you feel stuck and need a hint, and even then just trying to get the hint and as little extra as possible. If one reads a paper that way, then one is constantly solving problems, some just exercises and some quite a bit harder. It seems likely that an LLM could get to know what our mathematical background is and feed us with just the right hints to allow us to work our way through a mathematics paper in this active way. Yes, we would lose the particularly deep level of understanding and ownership that comes with having solved a hard problem oneself, but it isn’t clear to me that progress in mathematics would suffer as a result. I would be very interested to hear counterarguments to precisely this point. That is, I would be interested to know what use that level of deep involvement with a proof might have in a world where AI is much better than we are at finding proofs.

How about collective understanding? Let me quote a bit more of the letter.

Indeed, the mass production at faster and faster pace of “true/false” statements could destroy fertile ground instead of breathing life into new ideas.

Often these solutions are announced in a rush, leaving no time for a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others. As in all creative professions, this raises severe attribution and plagiarism questions. Moreover, without the willing mathematicians who must take care of their development and integration into the mathematical canon, AI-conceived ideas would never become fully alive and the crucial human transmission chain between mathematicians would be lost.

I’ll come back to questions about proper citation and focus on what I take as the core worry here: that if results are proved too quickly, then the digestion process will become impossible. I am definitely worried that results will not be properly digested, but for different reasons.

A first remark is that what AI is producing is not just true/false statements: we now know not just that the Navier-Stokes equation with smooth forcing admits finite-time blow-up, but we have a proof of that, which builds on a great deal of wonderful work done by human mathematicians. Many people used to express the worry that AI would solve our favourite problems with utterly opaque proofs, but that has not turned out to be the case, even if their write-ups often leave plenty to be desired. (Incidentally, I see these inadequate write-ups as almost certainly a temporary annoyance and therefore not as a fundamental threat to mathematical practice or future mathematical understanding.)

Secondly, even if the volume of new results is large, mathematics is a highly specialized discipline, so mathematicians can work in parallel. If, for example, we had to digest 1000 important results in a year that were roughly uniformly distributed across mathematics, then most sub-communities of mathematicians would probably want to understand around 30 of them, and for each individual problem there might well be only a small handful of specialists who would be obvious people to take the lead in reaching this understanding, with that handful varying from problem to problem. So it would be a big task, but not necessarily an impossible one.

In this context, it is worth thinking about the huge volume of output of human mathematicians, which seems to have been increasing recently, even before AI. While I have sometimes heard complaints about this, I have certainly not heard suggestions that human mathematicians should slow down the rate at which they prove interesting theorems. That may be partly because the authors of those theorems take the trouble to write their papers well and give good talks. But what about the large quantity of papers, including important ones, that are not written well and whose authors give incomprehensible talks? That can be annoying, but it is a familiar annoyance and not one that we think of as a crisis.

A third point is that even if the volume of AI output is too big for us to be able to digest it properly, that is not necessarily a bad thing. To draw an imperfect analogy, there is now more content available on streaming services than anyone could possibly watch, with the result that there is almost certainly some very good content out there that is hardly watched at all. But that isn’t obviously a worse situation than if there were far less content and all of it received the attention it deserved. Returning to mathematics, if there were too much AI-generated content for us to be able to digest it, then we could choose which parts of it we wanted to digest.

For that we would need to have some idea what was there (a situation a little similar to how human mathematicians typically learn quite a lot about what results are known in their area even when they do not understand their proofs in any detail). One way one could try to achieve that would be to create a well-designed database, probably with AI help. But perhaps that would be unnecessary, and instead one could simply talk to an LLM and ask it to give a bird’s-eye view of whatever area of mathematics one wanted to understand in that knowing-what’s-there way.

The fear seems to be that some very interesting and important parts of mathematics will be discovered by AI and then overlooked, when had they been discovered by human mathematicians they would not have been overlooked. And that may even be the case, but what matters is whether the amount of interesting and important mathematics discovered by AI that is not overlooked will exceed the amount of interesting and important mathematics that would have been discovered and properly digested by humans with AI having played a more modest role.

In short, it seems to me that while a flood of “big” AI results would be likely to increase the amount of important mathematics that was not properly digested, it would also be likely to increase the amount that was properly digested, which seems like a pretty good bargain.

Let me quickly discuss the problem of AI not properly crediting human mathematicians. I agree that this is a serious problem right now, but it is another problem that I see as temporary. Very soon, the whole “credit system” will surely collapse, since finding an amazing proof will be no more of an intellectual achievement than when a citizen scientist spots through their telescope an object that turns out to be a new comet. Until that happens, it is important to give humans the credit they deserve, since careers can depend on it, but that will soon cease to be the case as well. I have to say that I’m puzzled that this problem exists, since I would have thought that if you asked an LLM to look at a proof and tell you which ideas in it are close to ideas that are in the literature already, it would be extremely good at that task. I hope the answer to this conundrum is not that people have been in such a hurry that they have simply not taken the trouble to do this, but I fear that it might be, at least in some cases. If so, then those who have been careless deserve to be criticized, but it is a minor matter compared with the survival of mathematics, especially if the lack of citations is swiftly put right.

Does all this mean that I am optimistic that mathematicians will end up digesting at least as much mathematics in a post-AI world as it would have if AI had not been able to prove major theorems? Not exactly. But my worry is not that we would be unable to do it, but rather that the social structures that currently support this digestion process will be destroyed and not adequately replaced.

One way that might happen is that AI disrupts society so much, or even kills vast numbers of us, that the preservation of something like the current mathematical tradition ceases to be of any concern: all that will matter is the survival of the human race. But that again is a topic for a different blog post (which in fact I am in the middle of writing).

Let’s assume instead that we get lucky and that AI remains more or less under control. My worry then is that we do not manage to transmit what we know to a new generation of mathematicians. Speaking for myself, my main motivation for becoming a mathematician was the dream that I would solve unsolved problems — the more famous the better. I have also always greatly preferred directly thinking about a problem to reading books and papers and generally learning the mathematics of other people. (I’m not saying that’s good, but just stating a fact about myself.) If the dream of solving a famous problem had not existed, I’m not sure whether I would have become a mathematician. I don’t completely rule it out: maybe what really motivated me was that I had an aptitude for the subject and that solving problems was a way of getting respect from a small group of peers. And maybe I could have tried to gain that respect in a different way, such as thinking very hard about an area of mathematics until I was able to demonstrate to others just how well I understood it. But I’m not sure how motivating that would have been for me. I very much hope that there is a pool of young people for whom it will be a powerful motivation, because I think the survival of a human mathematical tradition may well depend on it.

Thus, the primary risk, as I see it, is that a lot of people who would have done a PhD in mathematics and gone on to become custodians of the mathematical tradition will no longer wish to do so. Those of us who have PhD students, including me, need to try as hard as we can to come up with imaginative ways for them to use their time productively (in consultation with the students themselves, obviously). Whether or not we do a good job with that could make a huge difference to the future of mathematics. A related risk is that the perception among policy-makers will be that mathematicians are no longer needed and that funding will become much harder to come by: we urgently need to come up with good ways of explaining the value of having a large pool of human mathematical experts, even if it is no longer part of their role to find new proofs of theorems.

A final reason that I didn’t sign the letter is that I wasn’t really sure what it was demanding that isn’t happening already. It seems likely that in a matter of not very many months LLMs will be released that are able to solve major mathematical problems, and they will presumably have no trouble at all with more run-of-the-mill problems. However much we might regret that, there is no chance that the impact of such models on mathematics will persuade AI companies to stop their release, though perhaps concerns about safety will lead to some delay and give us a bit more time to work out how to adapt. Assuming that they are released, there will be a flood of new results, whether we like it or not, and it will no longer be the AI companies producing them, though perhaps the pattern will continue that the AI companies will have access to more powerful models and so will obtain more than their fair share of headline results. So I felt that there was nothing to be gained from criticizing AI companies for generating too many solutions too quickly. In fact, it may well be that all that does is bring forward by a couple of months what was going to happen anyway, and perhaps it will even allow the results to be released in a more controlled way than they would have been if they had been discovered by random people once the models were publicly available. Under the circumstances, I think the best we can do is recognise the changes that are coming and try to work out the least unsatisfactory way of dealing with them.

Scott Aaronson Announcing BQP Partners: my and my brother’s new angel-investing venture

As I’ve written before, these past couple years I’ve often felt like the last remaining person in either quantum computing or AI who lacked a stake in some startup company whose valuation is right now shooting into interstellar space. My academic colleagues, including the ones who seemed the most singleminded about quantum oracle separations and other gloriously useless pursuits? One by one, like in a zombie movie, I learn that they too have now launched startups, and invariably raised tens of millions of dollars, for the sorts of ideas we might’ve idly traded at coffee breaks back in the day, before getting back to our real work.

So why didn’t I join this rollicking party? Partly because of a lifelong fear that, the instant my self-worth became tied to how much money I made, I’d need to humble myself before people who bluster and bully and lie and hype and conceal … yet who nevertheless succeed at becoming orders of magnitude richer than me. I’ve been terrified of even starting down that road, of whether I’d still be myself at the end of it.

It’s also partly that I can’t stand failure, or regret, or being wrong. Of course, as an academic researcher I also fail, and regret things, and am wrong constantly—but there it feels tolerable, because normally I can tell myself that it’s all just down to my inborn limitations. After all, if I could’ve solved the major open problem that someone else solved, or written the brilliant book that someone else wrote, then presumably I would’ve done it!

Clearly, though, I could’ve mined bitcoin in 2010. I could’ve gotten an early stake in Amazon or Google. It’s not even like those ideas never crossed my mind. I just … didn’t act on them, for some reason. (But even if I had, I’d probably just be full of regret that I hadn’t done even more.) Thus, my only way to avoid paralyzing regrets, has been to tell myself constantly that I’m not in the forecasting or money-making businesseses in the first place.

It helped that, insofar as I’m shallow or covetous, insofar as I’ve desired things of this world rather than insight or eternal truth, it’s never really been money that I cared about, but just being respected and liked. Elon Musk is the richest man on earth, but also one of the most despised—which isn’t a bargain that I could imagine ever appealing to me.

Plus, when I actually meet billionaires, I don’t find myself envious of their mansions or cars or anything else that they have; I don’t feel like such things would make my life any happier. Maybe I slightly envy their ability to fund the causes they care about, or their professional staffs who relieve them of drudgery, but mostly I envy the way their wealth announces, to whatever extent it does: “I was right when others weren’t.” Again, though, I’ve never trusted the world to cause me to be right about the future valuations of companies or anything similar, so I’ve settled for having been right about PostBQP and algebrization and BosonSampling.

The bottom line is that I made a choice decades ago to forgo trying to get rich, no matter how many of my friends did the same, and to strive instead to discover and tell the truth—to be a professor, a blogger, a jokester, and an “objective” arbiter and commentator. “Then, surely, everyone will like me!” my internal monologue went. “Then, surely, they’ll be grateful for all the free service I’ve rendered them—for decades of blogging, without once so much as asking for a donation or running an ad!”


HAHAHAHAHAHA.

As any regular reader will know, my attempts to be loved as a blogger backfired pretty spectacularly. Or rather: they did lead to thousands of strangers liking me (and I’m grateful for every last one of you), but they also led to probably an order of magnitude more strangers hating me, and congregating on Reddit and Twitter and elsewhere to discuss how badly I suck. And of course, trying to shift that balance by writing what people want to hear, rather than what I actually believe, was never within my realistic option set.

In the startup context, it didn’t matter how carefully I avoided taking a direct stake for or against any of the companies I blogged about. People on Twitter simply assumed that I had a stake—for example, that I must’ve shorted D-Wave or IonQ, or invested in their competitors, or had equity in AI companies. For why else would anyone write what I wrote?

Amusingly, my attackers here typically did have precisely the conflicts-of-interest that they falsely accused me of having, but that was never at issue; only my imaginary conflicts-of-interest were. Even as the Scott-haters greedily filled their pockets (or tried to), I alone needed to keep turning my pockets out to prove that they were still empty.


So then, screw it! In partnership with my brother David Aaronson, who’s long done investing professionally, and on David’s guidance and encouragement, I’m hereby embarking on a new policy.

Namely: when I hear about a brand-new startup that sounds relevant to my interests—in quantum, AI, or anything else—and I like and trust the founders (ideally, because of their previous academic research work), David and I will often make a small seed investment if the founders are open to it. Or, of course, we might become advisors or get involved in some other way.

In fact, David and I are launching BQP Partners—the link goes to our AngelList, where you can read about how to invest with us if you’re interested, if you’re an accredited investor. (See also whether you can spot any differences between David’s writing style and preoccupations and mine!)

So far, David and I are investing in:

I have little doubt that more potential investments will come our way very soon (some, probably, as a direct result of this post).

Crucially, I can handle my burden of regret—the “why didn’t I do this much earlier, if I was going to do it at all?” question—by telling myself that friends of mine were not founding companies left and right until very recently. I can also tell myself that I’m doing this less as a bet about the future (in which case … what if I’m wrong?), than simply as a way to support brilliant colleagues doing things that I genuinely admire.

When I blog about a company, I’ll always disclose if I have a financial position that presents a clear conflict of interest, so you can judge for yourself whether to listen to me. (Although, if that’s the sort of thing you’d demand, then you probably weren’t listening to me in the first place, were you?)

Having reflected on it a lot these past few months, I’m happy with my new policy and with my and David’s new venture, and I’m curious to see where it goes. I’m at peace with the possibility that we’ll lose our shirts, but I’m even at peace with a more disturbing possibility—that we’ll make millions and then people will scream at me online for being a sellout, a hack, and a shill. Those people, as I’ve learned, were going to scream at me anyway.

September 15, 2026

Tommaso Dorigo — Radiacode Zero - A Powerful New Radiation Detector

Radiacode Zero - A Powerful New Radiation Detector

Ionizing radiation is all around us. We do not notice it: we have not developed any sense to detect it. Yet it may affect us in very serious ways, particularly because its effect on living cells and organisms is cumulative: a progressive degradation.

Tommaso Dorigo
Categories

September 14, 2026

Jordan Ellenberg — I love working here

Tommaso Dorigo — On The Annihilation Risk From AI

On The Annihilation Risk From AI

The debate on the risk connected with the development of superintelligent systems has been going on for a while now, and in the last few years it has intensified considerably - especially since large language models have established themselves as powerful new oracles, mathematics superpowers, and code-writing wizards.

Tommaso Dorigo
Categories

September 13, 2026

Doug Natelson — Recent superconductivity results + open positions at Rice

Much as I feel like I should write about the latest developments in US science policy, instead I want to point out two exciting recent superconductivity results.  Below I will also append a couple of other items, including open positions at Rice.
  • After Fig. 2b from here
    In this paper, researchers demonstrated high temperature superconductivity in a monolayer of Bi\(_2\)Sr\(_2\)CuO\(_{6+\delta}\) (Bi-2201).  The monolayer contains just a single CuO\(_2\) plane, and remarkably, the superconducting transition is only suppressed about 10% from the bulk value of around 35 K.  The authors were able to explore the phase diagram by tuning the oxygen content in situ, using vacuum annealing to drive out oxygen and ozone exposure to (seemingly gently) put it back in.  This allows them to examine a large swath of temperature/doping/magnetic field parameter space, showing evidence of critical scaling of the resistance near the transition as well as an anomalous metallic state.  There's a lot to digest here.  The mapped out zero-field phase diagram in a single device (shown here) is extremely impressive.  Studies like this can hopefully give new insights into what physics is truly essential to achieve high temperature superconductivity.
  • In this paper, investigators placed exfoliated NbSe\(_2\) encapsulated by hBN in a split-ring resonator cavity, and they observed enhanced critical temperature (by 0.15 K out of 6.53 K, or an increase of 2.3%), critical field, and critical current when the resonance frequency of the cavity is such that it apparently couples to superconducting fluctuations in the material on the spatial scale of the cavity.  There is a ton of interest in using electromagnetic cavities to modify the properties of quantum materials - see this review.  As far as I know, this is the first time that coupling to the vacuum mode of a cavity has actually enhanced superconducting properties.  Exciting times.
It's worth noting that both of these papers come out of groups in China - Changgan Zeng at USTC and Yuanbo Zhang at Fudan.   

In other news:
  • The NSF is going to make about half the number of awards this year as it did in The Before Times (2021-2024), according to this news article in Nature.  Figure 1 (shown here) is striking.  The claim is that the NSF leadership is taking clawed-back FY26 funding of around $1B and saving it for some as-yet unspecified, unannounced OSTP "grand challenges" project.  
  • NSF also announced "new" funding opportunities here.  As described in that article linked above, these are not exactly new - it's essentially a reorganization/rebranding of much of the NSF's portfolio now that they've eliminated divisions and retired older funding solicitations.  Noteworthy is that the amount of funding mentioned in these solicitations is all considerably lower than what the aggregate of the older solicitations used to have.  As a non-expert, it looks a lot like these solicitations are being prepared as if the presidential budget requested funding levels (you know, the ones that want to cut NSF by more than half) are the baseline.
Meanwhile, at Rice we have some faculty searches underway:
  • The Rice Advanced Materials Institute is searching for an assistant professor with an expertise in computational materials (including AI/ML).  See here.
  • Our chemistry department is searching for an assistant professor position with an emphasis including physical chemistry.  See here. 
  • There will also be an AMO physics position posted shortly - I'll update with the link when that becomes available. Update:  See here.
Finally, Nano Letters is having a seed grant competition for grad students.  It's not much money, but it is good experience and can inspire graduate student creativity. (Full disclosure: I'm an associate editor for the journal.)


Doug Natelson — Science communication - importance, insights

This past week we launched SCOPE, a new center for science communication and public engagement.  We marked the occasion with a fun symposium, as well as a Science Café event the preceding evening and a public science openhouse yesterday.  The symposium was very enjoyable, with a panel that comprised Kelly Weinersmith (known for many things, including an outstanding podcast and popular science books such as this Hugo-award-winner), Peter Hotez (tireless champion of vaccine development and pushing back on disinformation), Eric Berger (space editor for Ars Technica, founder of spacecityweather and theeyewall, two excellent sites for no-hype weather information), and Briana Rapini (one of The Amoeba Sisters, creators of a youtube channel with 2.9M+ followers).

You might have picked up from my 21 years of blogging that I think science communication is of great importance.  We've learned amazing things about how the world works, and I think we'd all be better off if more people knew about them and about the process of learning and discovery.  If there is public investment in research, then it's incumbent upon researchers to make sure that the public has the opportunity to learn about the fruits of those labors.  When the government, NGOs, and corporations make policies and strategic decisions that involve or depend on technical knowledge, we need to do our best to help those be informed decisions.   Once upon a time, Congress had a research office to help their staff and office holders understand technological issues.  It was killed in 1995 as "wasteful" and allegedly partisan. <sarcasm> thank goodness no technology-oriented issues have come up before the US government since then.</sarcasm>  (I am very tired of victim-blaming that presents mistrust of science or partisan razing of the research ecosystem as somehow the direct fault of scientists who failed in the communication mission.  Communication could have been better about many things, but complex societal forces are, in fact, complex, and there are many deep-seated reasons behind where we are right now.)

There were a few key points that came out of the panel above and from related discussions at the symposium.

  • Know your audience and put yourself in their place.  What would you want to hear?
  • Respect your audience.  You can avoid jargon without condescension.
  • Ascribed to my colleague Neal Lane:  "The general public expects that you're smart.  They want to see if you're human."
  • Listen to your audience.  Ascribed to Will Rogers:  "Never pass up a chance to shut up."
  • If you're hesitant to do your public-facing project (writing, podcast, videos, etc.) because it's not flawless, just push through and do it.  The way to get good at this is through practice, not perfectionism.
There is a real dilemma out there about the degree to which practicing scientists can and should put effort into science communication.  Very few people in the US can name a single active scientist.  Among scientists and engineers, there is still sometimes an attitude of "Why are you spending your time on this?  If you are, you must not be a serious researcher."  It is true that, if you're a faculty member teaching and running a research program, you have to carve out time to do this, and those efforts are historically not well rewarded by many evaluation schemes.  Yet, I still think it's important, and programs like those run by SCOPE are hopefully going to help those who have an interest in science communication develop their skills and get valuable experiences.




September 11, 2026

Scott Aaronson 9/11 in Berkeley

Note: Of course I’ve been glued all week to the dramatic developments in AI. I’m working on a post about them. I’m not good at reacting to things in a timely way. So today, I’ll do my post marking the tragedy a quarter-century ago that we all commemorate. Please feel free to share your 9/11 memories in the comments. Also, Shana Tova to those who celebrate!


The morning of September 11, 2001, I was a second-year PhD student at Berkeley, who woke up late in his dorm room at International House, after a long night spent closing in on the proof of the quantum lower bound for finding collisions.

Rolling over to my laptop, I saw a flurry of weird emails, including one from Prof. Christos Papadimitriou saying that “we’re a community, and we’ll all support each other,” and another from Prof. Luca Trevisan (whose algorithms course I was then TA’ing) saying “on a day like this, it’s impossible to think about algorithms. Class is cancelled.”

Confused, I clicked over to the New York Times and saw the picture of the burning towers, and read numbly about what was already over by the time I’d woken up. I checked in with my mom, made sure relatives and friends in the NYC area were OK. My dad was at a company event in Atlanta, and would need to drive home because of the national grounding of flights.

One of my earliest memories in life, from age 5, is of ascending to the top of the World Trade Center. Growing up an hour’s drive from NYC, it wasn’t an exotic place to me.

I soon learned that one of the dead was Danny Lewin, the ex-IDF captain, theoretical computer scientist, and cofounder of Akamai who had his throat slashed on one of the planes while trying to fight the hijackers, making him the day’s first casualty, even while Akamai’s technology was part of what kept news websites running that day. I’d never met Danny but already knew many people in common with him. A few years later I’d be humbled to win the student paper award that was named in Danny’s memory.

Anyway, at Berkeley on 9/11, I wandered over to Soda Hall just to be with other people. A few students showed up for office hours, wanting help with their algorithms homework, which I found hard to believe, but I did my best to concentrate, as the computer screens around me showed the burning towers.

That evening, I went to a vigil for the victims in Sproul Plaza. But the “vigil,” such as it was, quickly dispensed with mourning and prayers and turned to applauded speeches about how the US must respond with love rather than war, and must turn the other cheek. Meanwhile, a student communist organization was handing out flyers explaining that the victims were mostly “wealthy capitalists and the workers who tried to rescue them.” This while smoke still blanketed NYC and the desperate search for survivors continued. I left the vigil early.

Until that day, I had thought of myself as basically a “leftist,” one whose #1 issue was the existential risk of climate change. Sure, I disagreed with my fellow leftists about issues from nuclear power to gifted education to Israel, but those were just intra-left disputes.

The year before, I had created the website “In Defense Of NaderTrading,” in a desperate attempt to intervene in history and cause Al Gore to become president rather than George W. Bush. When Bush “won,” by the infamous 537 votes in Florida, I considered it a victory for horribleness that would never be surpassed by anything else in my lifetime (ha). I couldn’t imagine any politician who was more the antithesis of everything I believed in than Bush. This view, of course, did not particularly stand out at Berkeley.

In the days after 9/11, though, it became obvious that I could not be a “leftist” in the Berkeley sense. Some of my fellow students felt that Osama bin Laden made a lot of great points, that the attacks were basically justified, and that at any rate, we in Amerikkka had done much worse to provoke them, including by supporting the genocidal settler-colony called “Israel,” which for all we know secretly masterminded the 9/11 attacks anyway (although again, if bin Laden had done them, he would’ve been justified).

Around the same time came the Second Intifada, when a wave of suicide bombings in Israeli buses and pizza parlors and university cafeterias thrilled and energized some Berkeley students to the extent that they took over a Holocaust Remembrance Day event with bullhorns to make it about the Nakba, smashed the windows of the Hillel building, and beat up a couple of students wearing kippot. That was how thoroughly anti-Nazi they were.

I finished my PhD at Berkeley in 2004 having learned about more than quantum computing. I’d learned that, while American academia had pockets that truly were crucial refuges and oases for nerds like me, it also harbored people who would gladly see me and my relatives and my fellow Americans killed for the sake of their ideological vision. And I’d learned that I had my own ideological vision, which was that such people could go fuck themselves.

It deeply pained me to be on the same side of anything as George W. Bush — especially because I knew that 9/11 had happened on his watch, that he had ignored all the warnings, and that he was grossly incompetent to manage the resulting wars against jihadism (just how incompetent, I didn’t know at the time). But as flawed as Bush was, I knew that I wanted to preserve rather than destroy the civilization of which he was a temporary steward. And I think the value and fragility of our civilization is the main lesson from that day that I’d like to convey to my kids, for whom of course 9/11 is just another historical event to learn about in school, like the Boston Tea Party or the Alamo.

Andrew Jaffe — The Talk I Gave on September 12

More 25th anniversary thoughts and recollections.

I had moved to Oxford less than two weeks before, still settling into my new life in the UK, working at Imperial College in London. That day, I was heading off to Durham, in the north of England, to a conference called “A New Era in Cosmology” — my first big talk now that I had started my permanent academic job. I was going to be a bit late, only able to arrive toward the end of the first day of the conference: September 11, 2001.

My train left in the late morning. A few hours in, passengers were starting to talk about an attack on New York City. This was in the days before smartphones and constant communication — I didn’t even have a mobile phone. The discussions around me were getting more and more frantic, and I was doing my best to piece together the story.

I was born in New York City, and much of my family still lived in the area. My parents lived in the suburb of Fort Lee, New Jersey, right across the Hudson River from Manhattan; from the apartment where I grew up, we had a fantastic view from the 18th floor. My father worked in The City, commuting every morning by car from New Jersey to midtown Manhattan. He would have been at his office that day.

I considered getting off the train somewhere en route, perhaps Sheffield or York, to try to get some more information, but I just stayed on the train. I was able to get a taxi from the station in Durham to my hotel, listening to the news. I was able to call my partner, back in Oxford, who had, luckily considering the pressure on transatlantic calls, been able to get in touch with my family in New York and New Jersey. Everyone was, thankfully, alright, although at this point my father was still in Manhattan. He had noticed all the emergency vehicles speeding downtown, thinking that it was a motorcade for some foreign dignitary, but it was the first fleet of emergency responders heading towards the twin towers after the first collision. Eventually, he made it home, to my family’s apartment overlooking the Hudson. But it was so close to the George Washington Bridge, a piece of vital infrastructure thought to be in danger after the first attack, that he had to be let off a mile or so away and walk the rest.

As for me, I had to give the first talk on 12 September, about the “new era in cosmology” that coming CMB measurements would usher in. We started, understandably, with a few moments of silence, and I remember trying to come up with some appropriate words with which to start, something about needing to persevere even in the face of terrible events. I also recall that it was given on old-fashioned transparencies (I should try to dig it out of my files…) and that it was actually one of the better talks I had ever given, calmed, or at least slowed compared to my usual nervous agitation, by the events.

(My friend and colleague Peter Coles was also at the conference, offering his own reminiscences on his blog.)

As I mentioned in my last post, 9/11 feels like a milestone, and a millstone — the world hasn’t been the same since. And we still need to persevere in the face of terrible events.

Andrew Jaffe — The Talk I Gave on September 12

More 25th anniversary thoughts and recollections.

I had moved to Oxford less than two weeks before, still settling into my new life in the UK, working at Imperial College in London. That day, I was heading off to Durham, in the north of England, to a conference called “A New Era in Cosmology” — my first big talk now that I had started my permanent academic job. I was going to be a bit late, only able to arrive toward the end of the first day of the conference: September 11, 2001.

My train left in the late morning. A few hours in, passengers were starting to talk about an attack on New York City. This was in the days before smartphones and constant communication — I didn’t even have a mobile phone. The discussions around me were getting more and more frantic, and I was doing my best to piece together the story.

I was born in New York City, and much of my family still lived in the area. My parents lived in the suburb of Fort Lee, New Jersey, right across the Hudson River from Manhattan; from the apartment where I grew up, we had a fantastic view from the 18th floor. My father worked in The City, commuting every morning by car from New Jersey to midtown Manhattan. He would have been at his office that day.

I considered getting off the train somewhere en route, perhaps Sheffield or York, to try to get some more information, but I just stayed on the train. I was able to get a taxi from the station in Durham to my hotel, listening to the news. I was able to call my partner, back in Oxford, who had, luckily considering the pressure on transatlantic calls, been able to get in touch with my family in New York and New Jersey. Everyone was, thankfully, alright, although at this point my father was still in Manhattan. He had noticed all the emergency vehicles speeding downtown, thinking that it was a motorcade for some foreign dignitary, but it was the first fleet of emergency responders heading towards the twin towers after the first collision. Eventually, he made it home, to my family’s apartment overlooking the Hudson. But it was so close to the George Washington Bridge, a piece of vital infrastructure thought to be in danger after the first attack, that he had to be let off a mile or so away and walk the rest.

As for me, I had to give the first talk on 12 September, about the “new era in cosmology” that coming CMB measurements would usher in. We started, understandably, with a few moments of silence, and I remember trying to come up with some appropriate words with which to start, something about needing to persevere even in the face of terrible events. I also recall that it was given on old-fashioned transparencies (I should try to dig it out of my files…) and that it was actually one of the better talks I had ever given, calmed, or at least slowed compared to my usual nervous agitation, by the events.

(My friend and colleague Peter Coles was also at the conference, offering his own reminiscences on his blog.)

As I mentioned in my last post, 9/11 feels like a milestone, and a millstone — the world hasn’t been the same since. And we still need to persevere in the face of terrible events.

Matt von Hippel — Everybody Who Isn’t ”Viewers Like You”

Last week, I talked about how truthseekers get paid. But truth-tellers and truth-seekers are different things.

Consider educational kids’ shows on public television.

Nobody who works on Sesame Street is out there uncovering new letters and numbers. Bill Nye’s show wasn’t bringing analysis fresh from the lab.

The purpose of these shows is to educate. The purpose of education is to change minds.

So who pays for educational kids’ shows on public television?

If you’re from the US and watched PBS growing up, you remember one answer: “viewers like you!” US public television is supported by donations, ordinary people across the country who want it to keep on educating kids.

But you also might remember the lists of names that came before “viewers like you”. Some of those were things like “the Department of Education” or “a grant from the National Science Foundation”: government programs, in other words. Others were philanthropists and private foundations. Some were tied to companies, like the Intel Foundation, or Juicy Juice.

All of these groups, from government departments to donors, are trying to change kids’ minds. They support specific shows on specific topics, where they want kids to be better-informed. The same groups have the same kind of impact on schools. For example, I remember in elementary school we all learned to play a recorder, because a wealthy donor had given the school recorders out of the idea that music education was especially important.

For a truth-seeker like a journalist, accepting that kind of funding would be a problem. Grants for journalists tend to support things like travel, letting journalists learn more about specific topics, not pre-judging the conclusion. But children’s television is about truth-telling, not truth-seeking, so our standards are different. We trust the people making children’s television to care about whether they’re telling the truth. And because the topics aren’t new, we don’t usually worry about their judgement being biased.

All this is rather obvious. But now, consider science YouTube.

Some science YouTubers seem to have a mission much like children’s television. They’re there to teach, not to make independent judgements. They don’t search for truth on their own. And some of them are funded by educational grants, much like children’s television.

Others are a bit more like journalists, or even activists. People follow them for their opinions, to hear their assessment. They’re trying to be truth-seekers.

On YouTube, it’s not always obvious which is which.

There’s a particular group of philanthropists called Effective Altruists, and many of them are concerned about AI. So in between funding things like anti-malaria bed nets, some of them are giving grants to YouTubers to make educational content about AI-related risks.

Apparently, they reached out to Sabine Hossenfelder, which was a bad idea. Sabine Hossenfelder’s followers aren’t just looking for education on known facts. They’re looking for her judgements, her literal bullshit-rating on ideas. And so while she’s paid by “viewers like you”, she’s not really the type to get paid by that type of grant.

What I want to emphasize, and what looked like it was getting lost in the discussion, was that their pitch would have been totally reasonable for other YouTubers. Educators do occasionally get grants to educate on specific topics. This is in fact a totally normal thing. Some YouTubers are educators first and foremost, they aren’t there as truth-seekers, but truth-tellers, with a real difference in how careful they need to be about bias.

Some YouTubers are different from other YouTubers. News at 11.

September 09, 2026

John Baez — The E6 Root Polytope

I’ve been thinking about the exceptional Lie algebra E6, as a spinoff of my project on E7, so I want to get a good mental picture of the E6 root polytope. This is 6-dimensional polytope with remarkable symmetry.

Let’s climb up to it, starting with some of its 4-dimensional faces, which are called 4-demicubes because you get them by taking a 4-dimensional cube, or tesseract, and removing every other corner. The 3-demicube is just a tetrahedron, since you can fit two tetrahedra in a 3-dimensional cube like this:

The 4-demicube builds on this fact in a surprising way.

I’m going to use the technology of Dynkin diagrams, or technically Coxeter diagrams: they’re closely related, and the difference is invisible here. I won’t explain them, just use them. I explained them here:

• Symmetry and the fourth dimension: part 3, part 4, part 5, part 6.

Let’s dive in!

The 4-demicube lives in 4 dimensions. It has 8 vertices.

You get it from a 4-dimensional cube, which has 24 = 16 vertices, by keeping every other vertex, throwing away half. That leaves 8.

What are its top-dimensional faces, aka ‘facets’? Surprise: there’s only one kind! All of them are regular tetrahedra.

In higher dimensions the demicube has two kinds of facet. You get a simplex-shaped facet from every other vertex, formed when you remove it. And you get a demicube-shaped facet from each of the cube’s facets. But in 4 dimensions the two kinds happen to be the same shape!

Eight of them are tetrahedra. These appear at the 8 corners you sliced off: one per removed corner.

Eight more come from the 8 faces of the 4-dimensional cube. These are 3-demicubes. But as we’ve seen, the 3-demicube is also a tetrahedron!

So the 4-demicube is especially symmetric: it has 16 tetrahedral facets. You can find coordinates where its vertices are

(±1, 0, 0, 0),   (0, ±1, 0, 0),   (0, 0, ±1, 0),   (0, 0, 0, ±1)

It’s actually one of the 4-dimensional regular polytopes, sometimes called the 4-orthoplex. It’s also called the 16-cell because it has 16 facets. It’s the 4-dimensional cousin of the octahedron, which has 8 triangular facets.

You can read some of these facts off the D4 Dynkin diagram, if you know what you’re doing. As you can see above, this diagram has a central node with three arms, each just 1 edge long: a perfectly symmetric three-pronged star. To get the 4-demicube, you ring the tip of any one arm.

To get the facets of the 4-demicube, delete an unringed node so the piece still holding the ring stays connected, and see what diagram survives. There are two choices: you can delete the tip of either other arm. But either way, what’s left is a straight chain of 3 nodes—the so-called A3 diagram—with a ring at one node at the end. This gives the tetrahedron.

Both choices give the same shape of facet, a tetrahedron, because all three arms of the D4 Dynkin diagram are interchangeable. That ceases to be true in higher dimensions!

 

Next, the 5-demicube. This lives in 5 dimensions and has 16 vertices.

You get it from a 5-dimensional cube—which has 25 = 32 vertices—by keeping every other vertex, throwing away half. That leaves 16.

What are its top-dimensional faces, or ‘facets’? There are two kinds!

Sixteen of them are 4-dimensional analogues of the regular tetrahedron, called 4-simplexes. These appear at the corners you sliced off: one per removed corner.

The other ten come from the ten faces of the 5-dimensional cube. After you take every other vertex, they become 4-demicubes. These are precisely the 4-demicubes we saw in the last section!

You can also read these two kinds of facets from the D5 Dynkin diagram. As you can see above, this diagram has three arms of lengths 2, 1, 1 (edges from the central branch node). To get the 5-demicube, you ring the tip of either length-1 arm. That ringed diagram encodes the whole polytope.

To get the facets, delete an unringed node so the piece still holding the ring stays connected, and see what diagram survives.

There are two choices.

If you delete the tip of the other length-1 arm, what’s left is a straight chain of 4 nodes—the diagram whose polytope is the 4-simplex. That gives the 4-simplex faces.

Or you can delete the tip of the length-2 arm. Then what’s left is a shorter branching diagram, the one I showed you in my last post! That gives the 4-demicube faces.

So the 5-demicube has both 4-simplex and 4-demicube faces.

Next let’s go up to the 6th dimension, which was my goal all along.

 

The E6 root polytope lives in 6 dimensions. It has 72 vertices.

What are its facets? You can read them straight off the E6 Dynkin diagram, using the same procedure we’ve been using so far.

As you can see, the E6 Dynkin diagram has three arms of lengths 2, 2, 1 (edges from the central branch node). To get the root polytope, you ring the node that’s the tip of a length-1 arm. That fact is not obvious, but let’s go ahead and do that.

Then, to get the facets, delete any unringed node such that the piece still holding the ring stays connected, and see what diagram survives.

There are two choices: the two other nodes at tips of the Dynkin diagram.

However, deleting either of these nodes leave a D5 diagram with a ring on one node, and this gives the 5-demicube we saw last time: a 5-cube with alternate vertices removed.

So the facets of the E6 root polytope are all the same shape: 5-demicubes!

With more work, we can count the facets of the polytopes we’ve been studying:

• The E6 root polytope has 54 facets, all 5-demicubes. They come in two kinds, because we had two choices of which node to delete, so there are really 27 ‘positive’ 5-demicube facets and 27 ‘negative’ 5-demicube facets.

• The 5-demicube has 16 4-simplex facets, one for each vertex that we removed from the 5-cube to create this demicube, and 10 4-demicube facets, one for each facet of that 5-cube.

• The 4-demicube has 8 3-simplex facets, one for each vertex that we removed from the 4-cube to create this demicube, and 8 3-demicube facets, one for each facet of that 4-cube. But both the 3-simplex and the 3-demicube are the familiar tetrahedron. So in fact the 4-demicube has 16 tetrahedral facets. Indeed, the 4-demicube is the 4-dimensional analogue of an octahedron: the so-called 4-orthoplex, or 16-cell.

Using some fancier math I explained here, we can count all the faces of the E6 root polytope. This polytope, is also called 122 due to the shape of its Dynkin diagram: the ring is on a branch of length 1, not counting the central node, while the other two branches have lengths 2. You can look up all this information on the Wikipedia page 122 polytope:

Faces of the E6 root polytope, or 122
dim faces count
5 5-demicubes 54 = 27 + 27
4 4-demicubes = 4-orthoplexes 270
4 4-simplexes 432 = 216 + 216
3 3-simplexes = 3-demicubes = tetrahedra 2160 = 1080 + 1080
2 2-simplexes = triangles 2160
1 1-simplexes = edges 720
0 0-simplexes = vertices 72

The 5-dimensional facets are all 5-demicubes, but as we’ve seen, they come in two kinds: that is, they lie in two orbits of the symmetry group. We can call 27 of them ‘positive’ 5-demicubes and 27 of them ‘negative’ demicubes. Of the 4-dimensional faces, 270 are 4-demicubes and 432 are 4-simplexes. Moreover the 4-simplexes come in two kinds: 216 are faces of positive 5-demicubes while 216 are faces of negative 5-demicubes. Let’s call the first kind of 4-simplex ‘positive’ and the second kind ‘negative’. The 3-dimensional faces are all tetrahedra, but they come in two ‘kinds’: 1080 of them are faces of positive 4-simplexes, and 1080 are faces of negative 4-simplexes. None is the face of both a positive and negative 4-simplex.

If you’re curious about how to count these things, see how some of us counted all the faces of the E8 root polytope here:

• John Baez, Integral octonions (part 5), The n-Category Café, September 3, 2013.

Here is a table of faces for the E7 root polytope, which is also called 231:

Faces of the E7 root polytope, or 231
dim faces count
6 221 polytopes 56
6 6-simplexes 576
5 5-orthoplexes 756
5 5-simplexes 4032
4 4-simplexes 16128 = 4032 + 12096
3 3-simplexes = tetrahedra 20160
2 2-simplexes = triangles 10080
1 1-simplexes = edges 2016
0 0-simplexes = vertices 126

Its 4-dimensional faces are all 4-simplexes, but they come in two ‘kinds’: that is, they lie in two orbits of the symmetry group of this polytope. Of the 4-simplexes, 4032 are the face of three 5-orthoplexes, while 12096 are the face of one 5-orthoplex and two 5-simplexes.

Here’s the E8 root polytope, also called 421:

Faces of the E8 root polytope, or 421
dim faces count
7 7-orthoplexes 2160
7 7-simplexes 17280
6 6-simplexes 207360 = 138240 + 69120
5 5-simplexes 483840
4 4-simplexes 483840
3 3-simplexes = tetrahedra 241920
2 2-simplexes = triangles 60480
1 1-simplexes = edges 6720
0 0-simplexes = vertices 240

There are two kinds of 6-simplex faces: 138240 of them each lie in one 7-simplex and one 7-orthoplex, while 69120 of them each lie in two 7-orthoplexes (and no 7-simplex).

Andrew Jaffe — Second test post

Andrew Jaffe — Test post

Checking some infrastructure…

Jordan Ellenberg — Wisconsin sports analytics and beer tomorrow night!

My colleage Sameer Deshpande, together with Shekhar Shah, and Paul Nguyen are doing Badgers on Tap Wednesday night 9/9 at 6:30pm at One Social Food Hall downtown; there will be talk about post-Moneyball sports analytics, beer, and trivia. Not sure I myself can make it but this is sure to be a good time with some savvy Badgers. Go!

September 08, 2026

n-Category Café The E6 Root Polytope

I’ve been thinking about the exceptional Lie algebra E6, as a spinoff of my project on E7, so I want to get a good mental picture of the E6 root polytope. This is 6-dimensional polytope with remarkable symmetry.

Let’s climb up to the E6 root polytope starting with some of its 4-dimensional faces, which are called 4-demicubes because you get them by taking a 4-dimensional cube, or tesseract, and removing every other corner. The 3-demicube is just a tetrahedron, since you can fit two tetrahedra in a 3-dimensional cube like this:

The 4-demicube builds on this fact in a surprising way.

I’m going to use the technology of Dynkin diagrams, or technically Coxeter diagrams: they’re closely related, and the difference is invisible here. I won’t explain them, just use them. I explained them here:

• Symmetry and the fourth dimension: part 3, part 4, part 5, part 6.

Let’s dive in!

The 4-demicube lives in 4 dimensions. It has 8 vertices.

You get it from a 4-dimensional cube, which has 24 = 16 vertices, by keeping every other vertex, throwing away half. That leaves 8.

What are its top-dimensional faces, aka ‘facets’? Surprise: there’s only one kind! All of them are regular tetrahedra.

In higher dimensions the demicube has two kinds of facet. You get a simplex-shaped facet from every other vertex, formed when you remove it. And you get a demicube-shaped facet from each of the cube’s facets. But in 4 dimensions the two kinds happen to be the same shape!

Eight of them are tetrahedra. These appear at the 8 corners you sliced off: one per removed corner.

Eight more come from the 8 faces of the 4-dimensional cube. These are 3-demicubes. But as we’ve seen, the 3-demicube is also a tetrahedron!

So the 4-demicube is especially symmetric: it has 16 tetrahedral facets. You can find coordinates where its vertices are

(±1,0,0,0),(0,±1,0,0),(0,0,±1,0),(0,0,0,±1)(\pm 1, 0, 0, 0), \quad (0, \pm 1, 0, 0), \quad (0, 0, \pm 1, 0), \quad (0, 0, 0, \pm 1)

It’s actually one of the 4-dimensional regular polytopes, sometimes called the 4-orthoplex. It’s also called the 16-cell because it has 16 facets. It’s the 4-dimensional cousin of the octahedron, which has 8 triangular facets.

You can read some of these facts off the D4 Dynkin diagram, if you know what you’re doing. As you can see above, this diagram has a central node with three arms, each just 1 edge long: a perfectly symmetric three-pronged star. To get the 4-demicube, you ring the tip of any one arm.

To get the facets of the 4-demicube, delete an unringed node so the piece still holding the ring stays connected, and see what diagram survives. There are two choices: you can delete the tip of either other arm. But either way, what’s left is a straight chain of 3 nodes—the so-called A3 diagram—with a ring at one node at the end. This gives the tetrahedron.

Both choices give the same shape of facet, a tetrahedron, because all three arms of the D4 Dynkin diagram are interchangeable. That ceases to be true in higher dimensions!

 

Next, the 5-demicube. This lives in 5 dimensions and has 16 vertices.

You get it from a 5-dimensional cube—which has 25 = 32 vertices—by keeping every other vertex, throwing away half. That leaves 16.

What are its top-dimensional faces, or ‘facets’? There are two kinds!

Sixteen of them are 4-dimensional analogues of the regular tetrahedron, called 4-simplexes. These appear at the corners you sliced off: one per removed corner.

The other ten come from the ten faces of the 5-dimensional cube. After you take every other vertex, they become 4-demicubes. These are precisely the 4-demicubes we saw in the last section!

You can also read these two kinds of facets from the D5 Dynkin diagram. As you can see above, this diagram has three arms of lengths 2, 1, 1 (edges from the central branch node). To get the 5-demicube, you ring the tip of either length-1 arm. That ringed diagram encodes the whole polytope.

To get the facets, delete an unringed node so the piece still holding the ring stays connected, and see what diagram survives.

There are two choices.

If you delete the tip of the other length-1 arm, what’s left is a straight chain of 4 nodes—the diagram whose polytope is the 4-simplex. That gives the 4-simplex faces.

Or you can delete the tip of the length-2 arm. Then what’s left is a shorter branching diagram, the one I showed you in my last post! That gives the 4-demicube faces.

So the 5-demicube has both 4-simplex and 4-demicube faces.

Next let’s go up to the 6th dimension, which was my goal all along.

 

The E6 root polytope lives in 6 dimensions. It has 72 vertices.

What are its facets? You can read them straight off the E6 Dynkin diagram, using the same procedure we’ve been using so far.

As you can see, the E6 Dynkin diagram has three arms of lengths 2, 2, 1 (edges from the central branch node). To get the root polytope, you ring the node that’s the tip of a length-1 arm. That fact is not obvious, but let’s go ahead and do that.

Then, to get the facets, delete any unringed node such that the piece still holding the ring stays connected, and see what diagram survives.

There are two choices: the two other nodes at tips of the Dynkin diagram.

However, deleting either of these nodes leave a D5 diagram with a ring on one node, and this gives the 5-demicube we saw last time: a 5-cube with alternate vertices removed.

So the facets of the E6 root polytope are all the same shape: 5-demicubes!

With more work, we can count the facets of the polytopes we’ve been studying:

• The E6 root polytope has 54 facets, all 5-demicubes. They come in two kinds, because we had two choices of which node to delete, so there are really 27 ‘positive’ 5-demicube facets and 27 ‘negative’ 5-demicube facets.

• The 5-demicube has 16 4-simplex facets, one for each vertex that we removed from the 5-cube to create this demicube, and 10 4-demicube facets, one for each facet of that 5-cube.

• The 4-demicube has 8 3-simplex facets, one for each vertex that we removed from the 4-cube to create this demicube, and 8 3-demicube facets, one for each facet of that 4-cube. But both the 3-simplex and the 3-demicube are the familiar tetrahedron. So in fact the 4-demicube has 16 tetrahedral facets. Indeed, the 4-demicube is the 4-dimensional analogue of an octahedron: the so-called 4-orthoplex, or 16-cell.

Using some fancier math I explained here, we can count all the faces of the E6 root polytope:

dim faces count
5 5-demicubes 54 = 27 + 27
4 4-demicubes = 4-orthoplexes 270
4 4-simplexes 432
3 3-demicubes = 3-simplexes = tetrahedra 2160 = 1080 + 1080
2 2-simplexes = triangles 2160
1 1-simplexes = edges 720
0 0-simplexes = vertices 72

If you’re curious about how to count these things, see how some of us counted all the faces of the E8 root polytope here:

Andrew Jaffe — 25 & 60

Twenty-five years ago, I moved from San Francisco to the UK, from a fellowship at Berkeley to a permanent job at Imperial College, London. A lot has changed since then. I was 35; now I am 60. It was two weeks before 9/11; the world hasn’t seemed as open and free since.

I arrived from the Bay Area just after the first dot-com bubble burst. London, adjusting to New Labour after almost two decades of Thatcher and Thatcherism, felt exciting and vibrant. But just weeks after I arrived came the horror of 9/11, and its years-long aftermath, especially the Iraq war which eventually doomed Blair’s premiership and probably led the way to the 2010 election, the disaster of “austerity” as a wrong-headed attempt to deal with the 2008 recession, and eventually to the more-disastrous Brexit on this side of the Atlantic. Similar politics, though with very different timing, led back in the USA to Obama, one of the few rays of political hope over the last quarter-century, but then, of course, to Trump. And everywhere since 2001 the rise of nativist populism making me feel at home, well, pretty much nowhere — a rootless cosmopolitan. Higher education, scientific funding, and curiosity-driven research are in a parlous state in both the US and the UK.

But: despite a few difficult years in the mid-2000s, I have prospered. Our analysis of data from the Planck satellite has solidified our standard cosmological model — but also given us new problems and puzzles to think and worry about. I have written a book, The Random Universe, trying to explain how we know what we know as scientists and as human beings. And my family, my wife and two daughters, are a source of joy and excitement that inspire me every day.

So now I am 60. I was honoured and humbled a couple of months to ago to be joined by many of my colleagues and scientific friends at a conference here in London. That, and getting a Transport For London 60+ Travel Card, makes it hard to avoid feeling old. But those colleagues and friends (many of whom are older than me) reassured me that it’s only the beginning of a new chapter.

Jordan Ellenberg — Finite-time blowup

Interesting developments tonight, as Levent Alpöge and Tristan Buckmaster announce that after a fair amount of work they have constructed examples of finite-time blowup for a broad class of PDEs including 3-d incompressible Euler, inspired by of Diego Córdoba and Luis Martínez-Zoroa, and using plenty of LLM iteration in order to get the details right. This is, of course, a problem in the neighborhood of Navier-Stokes (in the negative direction of finding a counterexample to the conjecture, which I have over the years heard many PDE folks saying was the right way to bet), and Terry Tao says in a Mastodon thread that in principle this method doesn’t seem so far from showing blowup for Navier-Stokes too, though a large amount of compute and detail-checking would be involved.

At least part of this has already been Lean-formalized, though perhaps eccentrically I find I care a little less about that. What matters is not whether there’s an example but whether the example has something to teach us. An interesting but incorrect example would surely be of more value than an uninteresting but correct one. Well, I suppose the latter would have more financial value. Though even on that Millennium Prize page, one sees: “Why ask for a proof? Because a proof gives not only certitude, but also understanding.” Very true! We mustn’t settle for mere certitude. Certainly the work of Alpöge, Buckmaster, Córdoba, and Martínez-Zoroa seems to offer understanding as well.

September 06, 2026

Jordan Ellenberg — Don’t Be Too Sure dramatis personae

I’m well underway on revising Don’t Be Too Sure, which I finished a first draft of right before surgery. A lot of people make appearances in this book, most of all William James and John von Neumann, who became the two main characters despite not being in my original plans for the book at all. Some other people: Felix Hausdorff, Alfred Kroeber and his daughter Ursula Kroeber Le Guin, John Keats, Sheila Heti, Katharine Briggs and her daughter Isabel Briggs Myers, Jakob Bernoulli, Elbert Hubbard, Anna Kiesenhofer, Grace Hopper, Thomas Jefferson, Caroline Hoxby, David Hilbert, Michel Adanson… well, there are a lot of people in it, who do a lot of things.

John Baez — The Mantle

As we descend from the base of Earth’s crust through the mantle, the rock does not remain unchanged. Pressure and temperature rise inexorably, and the minerals that thrive at the surface are forced, step by step, into new and denser crystallographic arrangements. This is the story of those transformations.

In this tale, I’ll act like I know a bit about minerals. I actually don’t: there are a bewildering variety, and I can never remember them. So don’t worry: when you come across a jargon-filled patch of prose, just power through it. You might learn a little… or you can just ignore it. The overall point here is that the Earth is made of beautiful crystalline structures that change character in complex ways as we descend.

The Mohorovičić discontinuity

Our story begins at the boundary where Earth’s crust, rich in feldspar and quartz, gives way to the denser mantle beneath. We see this boundary through its effect on seismic waves, and it’s called the Mohorovičić discontinuity or “Moho”. The Moho does not lie at one fixed depth: it’s 5–10 kilometers below the seafloor, but 30–50 kilometers below most continents, and as much as 70–80 below young mountain belts like the Himalayas.

The mantle just below the Moho mainly consists of a rock called peridotite, which is made mostly of olivine and pyroxene, with smaller amounts of garnet (or, at shallower depths, spinel). Peridotite has a delicious coarse green appearance:



More precisely, this is what peridotite looks like up here. But when geochemists talk about the bulk composition of the upper mantle, they often use an idealized model called pyrolite—not a rock you can pick up, but a hypothetical recipe Ted Ringwood proposed in the 1960s for the primitive upper mantle.

Why? Since the Earth has had a convecting mantle, solid mantle rock wells up in places. As it does, the pressure drops, and a bit of it melts: the minerals with lower melting points. This melt flows upward. It’s called basalt. It builds the Earth’s crust. But it leaves a residue behind, made of minerals with higher melting points.

In Ringwood’s theory, which for expository purposes I’ll assume is true, pyrolite is what mantle rock is like before any partial melting depletes it of basaltic ingredients. The name is a portmanteau of pyroxene and olivine, the two dominant minerals. Pyrolite is about 60% olivine; the remaining 40% is mostly pyroxenes plus garnet.

• A pyroxene is a mineral built from single, unbranched chains of corner-sharing SiO₄ tetrahedra, with metal cations—chiefly Mg, Fe, and Ca—linking the chains together. The general formula is XY(Si,Al)₂O₆, where X and Y are those cations.


• Olivine is a green silicate, (Mg,Fe)₂SiO₄:


Its crystal structure in the upper mantle is an orthorhombic arrangement of isolated SiO₄ tetrahedra knit together by magnesium and iron in octahedral sites. It’s called the α-phase because we’ll see some more compressed phases as we descend.

• A garnet is built from separate SiO₄ tetrahedra held together by cations, but assembled into a dense, hard, characteristically cubic-symmetry crystal. There are different kinds of garnet, but the general formula is X₃Y₂(SiO₄)₃: three divalent X cations, two trivalent Y cations, and three isolated silica tetrahedra. The mantle’s garnet is largely pyrope, Mg₃Al₂(SiO₄)₃.


As we descend, the pyroxenes and garnet gradually dissolve into each other, producing a new high-pressure mineral called majorite. Here’s a rare sample from a meteorite fall in Canada:


So even before the dramatic change 410 kilometers down, the rock is no longer the simple olivine-pyroxene-garnet assemblage we had further up.

The 410-kilometer discontinuity

Roughly 410 kilometers down, the pressure reaches about 13,000 atmospheres and the temperature hovers around 1,400°C. Olivine can no longer hold its familiar shape. It transforms to its β form: wadsleyite, a mineral with the same chemical formula but a fundamentally different atomic arrangement. Instead of isolated SiO₄ tetrahedra, wadsleyite contains paired Si₂O₇ groups, and the oxygens pack more densely. The density jump is sharp enough to be detected globally by seismologists as a reflector of earthquake waves.

Wadsleyite has a remarkable property: it can hold several weight percent of water locked within its crystal structure. The transition zone may thus contain more water than all the oceans combined! However, very little wadsleyite has been seen on the Earth’s surface. Here’s a bit from that same meteor fall in Canada:


The 520-kilometer discontinuity

Descend further, to around 520 kilometers, and the temperature goes up only a little, to roughly 1500–1600°C, since convection here is strong. The pressure goes up to about 175,000 atmospheres. At this point wadsleyite transforms into the γ form of olivine: ringwoodite. This is denser, still chemically Mg₂SiO₄, but now with cations packed into tetrahedral and octahedral holes in a close-packed oxygen framework—the most efficient packing geometry that nature offers for this composition:


Ringwoodite is named for the great Australian geochemist Ted Ringwood, who studied these transitions. Here’s an artificially manufactured sample:


For a long time the mineral’s existence in the mantle was purely hypothetical. But in 2014, a tiny grain was discovered as an inclusion inside a diamond brought up from the deep mantle by an eruption, providing the first direct proof of its existence in Earth’s interior.

The 660-kilometer discontinuity

At a depth of 660 kilometers and a pressure of roughly 230,000 atmospheres, the most dramatic phase transition of all occurs. Ringwoodite does not merely rearrange into a still more dense form! Instead, it decomposes into two entirely new minerals: bridgmanite (MgSiO₃) and ferropericlase (MgO). The majorite garnet also decomposes, yielding davemaoite (CaSiO₃), which is stable through the rest of the lower mantle:



The 660-kilometer discontinuity is sharp, globally consistent, and marks the conventional boundary between the upper and lower mantle. One reason it’s important is that enormous slabs of colder, denser rock sink through the upper mantle until they hit this discontinuity, where the phase change between ringwoodite and bridgmanite creates a kind of barrier.

These slabs are 30–100 kilometers thick and hundreds to a thousand kilometers across! Some punch straight through into the lower mantle and keep sinking. But many flatten out when they hit the barrier, sometimes lying there and piling up for tens of millions of years. You can see this in seismic images beneath Japan and the Marianas. Numerical models suggest that they pile up until they overwhelm the barrier and flush down in a comparatively sudden avalanche—lasting mere millions of years.

The lower mantle

This is the realm of bridgmanite, probably the most abundant mineral in the Earth. Bridgmanite is a beautifully symmetric cage of corner-sharing SiO₆ octahedra, with Mg tucked into the large cavities between them. It accommodates enormous pressure because there is very little void space left to compress.



It is a striking fact that while bridgmanite is the most abundant mineral on the planet, it went unnamed until 2014, simply because no natural hand-sized specimen had ever been recovered. Everything we know about it comes either from high-pressure laboratory synthesis, from microscopic grains in shocked meteorites, or from the indirect testimony of earthquake waves that have traveled through 2,000 kilometers of it.

For over 2,000 kilometers of descent, from 660 to roughly 2,700 kilometers down, bridgmanite and its companion ferropericlase reign without significant further phase change. Seismic velocities increase steadily, but there are no dramatic discontinuities.

The D″ discontinuity

As we approach the core-mantle boundary—at depths around 2,700 kilometers, pressures of approximately 120,000–125,000 atmospheres, and temperatures of 2,200–3,7000°C—even bridgmanite yields. It transforms into the post-perovskite phase. Post-perovskite is a layered, sheet-like structure of SiO₆ octahedra, quite different from bridgmanite’s three-dimensional cage, making it potentially much weaker and more prone to flow.

This transition is believed to be responsible for the seismic D″ discontinuity observed at 2,900 kilometers depth. The D″ layer is a highly dynamic region, likely the site of storage of subducted materials and the source of deep mantle plumes.

A summary of the descent

The table below summarizes the major transitions:

Depth (km)        Minerals
0–410 olivine (α) + pyroxenes + garnet
410 → wadsleyite (β)
520 → ringwoodite (γ)
660 → bridgmanite + ferropericlase + davemaoite
660–2700 bridgmanite dominates
~2700 → post-perovskite
2900 → liquid iron core

The interesting thing about this story is that it was told first by seismology—the sharp jumps in wave speeds at 410 and 660 kilometers were detected long before geologists could reproduce those pressures in the lab—and only later checked by diamond-anvil cell experiments squeezing tiny mineral samples to millions of atmospheres. The rocks never rise to the surface to tell their story directly, so much of the tale above is just theory.

Which minerals are there the most of?

We can estimate how much of the Earth is made of wadsleyite, ringwoodite, and bridgmanite using known shell volumes, estimated densities, and mineral proportions from the pyrolite model.

Step 1: Earth’s mass budget by layer

The Earth’s total mass is M⊕ ≈ 5.972 × 1024 kg. The mass budget by layer is approximately:

•    Crust: ~0.4% of Earth’s mass
•    Upper mantle + transition zone (35–660 km): ~18% of Earth’s mass
•    Lower mantle (660–2,891 km): ~49% of Earth’s mass
•    Core (outer + inner): ~32.5% of Earth’s mass

Step 2: The transition zone (410–660 km)

Using PREM densities averaging ~3,760 kg/m3 across the transition zone, and the volume of each spherical shell:

Wadsleyite zone (410–520 km):
Shell volume ≈ 4.8 × 1019 m3
Shell mass ≈ 1.76 × 1023 kg
Fraction of Earth’s mass ≈ 2.9%

Ringwoodite zone (520–660 km):
Shell volume ≈ 5.9 × 1019 m3
Shell mass ≈ 2.24 × 1023 kg
Fraction of Earth’s mass ≈ 3.8%

In the pyrolite model of mantle composition, forms of olivine (wadsleyite and ringwoodite) make up roughly 60% of the transition zone by mass, with the remaining ~40% being majoritic garnet. Applying this correction:

Wadsleyite: 0.60 × 2.9% ≈ 1.8% of Earth’s mass
Ringwoodite: 0.60 × 3.8% ≈ 2.3% of Earth’s mass

These estimates carry roughly 20–30% uncertainty, mainly from the assumed 60% olivine proportion in the transition zone, which varies with local temperature and bulk composition.

Step 3: Bridgmanite (660–2,700 km)

The lower mantle holds about 49% of Earth’s mass—it is an enormous shell! Bridgmanite constitutes approximately 80% of the lower mantle mineral assemblage (by mass) in the pyrolite model:

0.80 × 49% ≈ 39% of Earth’s mass

This is consistent with the well-cited literature figure that bridgmanite comprises approximately 38% of the planet’s mass—making it the single most abundant mineral in the Earth by a vast margin.

Mineral Depth (km) Fraction of Earth’s Mass
Wadsleyite 410–520 ~1.8%
Ringwoodite 520–660 ~2.3%
Bridgmanite 660–2,700 ~38–39%
All three combined 410–2,700 ~42%

Thus, these three minerals—all members of the same Mg₂SiO₄/MgSiO₃ chemical lineage—together constitute roughly 42% of Earth’s entire mass. All other named minerals on Earth, including quartz, feldspar, calcite, diamond, and the roughly 3,800 others known to mineralogists, divide up the remaining scraps.

John Preskill — How can objects interact without touching?

Rethinking the electric field

Have you ever wondered what an electric field actually is? 

The electric field is the foundation of most technologies that we rely on every day. From power grids and electronic devices to radio communication and the internet, the electric field is extremely relevant to our daily lives. However, despite its importance, I have always felt that the common explanations of the electric field leave something unanswered. 

Most textbooks define the electric field as a property of space or a physical entity surrounding electric charges, or with the equation of force per unit charge. These definitions help us understand what the electric field does and its effect on electrically charged particles, but they do not fully answer what an electric field actually is and how it influences charges. Thus, I started thinking about the question: what allows charges to influence each other without touching?

This question led me down a path that began with a simple observation in everyday life, and it eventually pointed toward much deeper ideas in modern physics.

Objects Influenced by Their Surroundings

Before talking about electric fields, let’s consider a more basic question: Does it seem reasonable that objects can be influenced by their surroundings? 

Most people would answer yes. We have all seen examples of objects responding to something else nearby, such as the Earth orbiting the Sun, a compass needle reacting to a magnet, and our phones responding to signals from a WiFi router. But what is the mechanism behind these interactions? 

A simple physical phenomenon that we can look at is a balloon rubbed on a piece of clothing that can pick up strands of our hair. Many of us have seen this demonstration in kindergarten or first grade of elementary school. This might seem completely ordinary, but if we pause and think about it, something strange is happening – the balloon is influencing the hair without touching it. 

How is that possible? One answer is simply that the balloon “pulls” on the hair, but this raises more questions. How does the balloon reach the hair? What is happening in the space between them? These questions suggest that something is missing from the picture of objects pulling on each other directly through contact.

Image of cat fur sticking to a balloon. Source: https://science.howstuffworks.com/why-do-balloons-stick-to-hair.htm

To put this in the context of physics, we might all have learned that “like charges repel, and opposite charges attract”. We might have solved equations on how fast charges would move away from or toward each other. We were always told to just accept it because these motions result from the electric field. But why do these observations happen? What is happening between the charges?

Historically, physics encountered the same problem. If one object can influence another at a distance, it is natural to ask what is happening in the space between them. One guiding principle that physicists often use is the concept of locality. Locality is the idea that an object can only be directly influenced by its immediate surroundings. Thus, an influence should not simply leap across space from one object to another, and changes should propagate through intermediate regions step by step. 

At first glance, locality seems reasonable because it matches many of our everyday experiences. If I push a book across a table, my hand influences the book through direct contact. The influence does not appear to jump instantaneously across the table. 

However, locality creates a tension when we return to the scenario of the balloon pulling on strands of hair. If locality is true, something must be happening in the space between the balloon and the hair. But from a standard electromagnetic perspective, the space between the balloon and the hair is empty. Therefore, we have encountered a contradiction: if what is between the balloon and the hair is empty space, then what is responsible for transmitting the influence? Neither the usual electromagnetism nor locality tells us the answer to these questions.

The Classical Electric Field

In the usual electromagnetic picture, the answer to the puzzle is the electric field. Rather than allowing charges to influence one another directly across space, the theory assigns an electric field to the space surrounding charges. The field acts as the intermediary through which influence is transmitted. 

A useful way to think about the electric field is that it assigns information to every point in space. If we imagine a charged particle that is placed at a particular location, the electric field tells us how that particle would move. This charged particle is what physicists call a test charge. By observing how the test charge behaves, we can infer information about the electric field at that location. 

This could seem like a satisfying answer as the electric field tells us how influence is transmitted, but it does not tell us what kind of thing is doing the transmission. Is the electric field a physical substance? Is it a mathematical tool? Or is it something else? 

It might be easy to fall back on the idea that the electric field ultimately works through tiny particles physically touching one another. After all, contact interactions are among the most familiar interactions that we experience. 

But physics challenges this intuition as well. It is surprisingly difficult to define what it means for two objects to “touch”. We usually think of the balloon attracting hair as an example of action at a distance, whereas pressing a hand on a table feels like direct physical contact. However, at the microscopic level, the two situations are fundamentally similar. If we could zoom in on our fingertip and the table with a microscope, we would find that the atoms in our skin never make contact with the atoms in the table. This is because of the repulsion between the electron clouds surrounding the atoms, which prevents the two atomic nuclei from overlapping. Say if we scale the atom in the table up to be the size of a marble, then the nearest atom in our fingertip would still be separated from it by a few centimeters. In the end, nothing is truly “touching” in the intuitive, physical sense. 

Thus, the idea of contact does not solve our problem. We are forced to ask the same question again: what is it that allows these interactions to occur? To answer that question, I turned to a different perspective of thinking about electric fields.

The Electric Field as A Dynamical Structure

From our intuition, it is natural to imagine the electric field as some invisible substance filling space. This is often the picture suggested by the common field line diagrams in physics textbooks, which make the field appear to flow outward or inward from charges, almost like a moving fluid. 

A useful analogy is the ocean. A boat floating on water can move because waves pass beneath it. The boat responds to changes in its surrounding waves rather than to some direct push from a distant object. Similarly, charged particles respond to changes in the electric field around them. We can then view the electric field as a dynamical structure that governs how the state of the world can evolve.

However, the ocean analogy can only take us so far. Ocean waves are made of water molecules. Sound waves are made of vibrating air molecules. But what is the electric field made of? When light travels through empty space, it seems that there is no material medium at all. 

This brings us back to the mystery: if locality suggests that something must exist in the space between interacting objects, and if the electric field is not made of the ordinary matter that we understand, then what exactly is occupying the space? 

To answer this question, we have to rethink what we mean by “empty” space itself.

Empty Space is Not Empty

Conventionally, we have always imagined empty space as exactly what the name suggests—empty. Just like if all the particles were removed and nothing was remaining. But modern physics suggests a very different picture. 

In quantum field theory, what we call “empty space” is not truly empty. Empty space is filled with underlying quantum fields that permeate all of space and time, even in the absence of particles. Even when the surface of the ocean looks perfectly still, the water is still there. The ocean is not defined only by visible waves, but by the underlying medium that can support waves in the first place. The waves are patterns of motion of the ocean itself, just like the electric field. These fields are part of the fundamental structure of the universe from which physical phenomena emerge. Quantum field theory suggests that particles are not independent objects moving through an otherwise empty space. Rather, they are localized patterns or excitations of underlying fields that already exist throughout the universe. 

From this perspective, the electric field is not something that is added to empty space. It is part of the fundamental dynamical structure of space itself.

Conclusion

At the beginning, I asked a simple question: how can objects influence each other without touching? The straightforward answer is the electric field. Charges create electric fields, and those fields determine how other charges move. But what is an electric field? Is it an invisible material filling space between objects, or is it a dynamical structure that governs how physical systems in the world evolve? 

From the perspective of quantum field theory, quantum fields permeate all of space and time. Particles are not separate objects moving through an empty space, but are excitations of these underlying fields; electric fields are not secondary matter surrounding charged particles, but are particular configurations of the underlying electromagnetic quantum fields. What we observe as the motion of a charged particle is the result of its interaction with the electromagnetic field, whose local state determines how the particle evolves. 

In the end, our original question may not have a single definitive answer. But asking it revealed a shift in perspective, and physics has repeatedly shown that every explanation opens the door to an even more fundamental question. Stopping at this step, a new mystery emerges: what are these underlying quantum fields themselves? What are they made of, and where do they arise from? 

September 04, 2026

Matt von Hippel — Paying the Truthseekers

Academics and journalists have a lot in common, at least in principle.

Whether you’re a reporter or a professor, your job is to go out into the world and figure out the truth. You’re supposed to be careful, to check and correct for how you might be wrong. And at the end of the day you’re supposed to communicate what you found.

The differences mostly come in how you’re paid.

You could imagine some sort of pure truthseeker, paid purely by how well they tell the truth. People would ask them to find out the truth about something, and pay them for the service. And the truthseekers with the best track record would get the most clients. But neither profession really works like this.

Journalism comes closest. Once upon a time, people bought newspapers in order to be the first to know when something important happened. While there’s still a little bit of that going on (I guess this is what Bloomberg Terminals are for?), it’s a lot less central because of the internet. Now, there are hundreds of ways to find out about things, from a multitude of news sites to social media. More and more, people expect to be able to get information for free.

In that environment, the news has to compete not on the facts themselves, but on how it presents them. People pay for news that’s curated well to match their interests, or news that feels more respectable. And more than either of those, they pay for news that’s entertaining. So while truthseeking skills pay, writing skills often end up mattering more. In a sense, it’s why it’s possible for me to do journalism at all. I was trained in the academic truthseeking tradition, not the journalistic one. I got into journalism by impressing editors with my writing, not my ability to suss out the truth.

That academic truthseeking tradition is quite different, in part because the rewards for it are much more indirect. Academics pay comes from two main sources: research grants, and student tuition. Students are mostly there to learn old facts, not new ones, so that source of money supports research only in so far that students believe that a successful researcher with time for research will also be a better teacher.

Research grants, in principle, pay for truthseeking. But they’re typically paid by governments, which often don’t have a clear idea of what they’d like to learn, since the more practical questions are already being researched by private companies. So the decision gets delegated out to other academics, who have a vague shared sense of what’s worth knowing and what’s not. Accuracy should have an impact: that is, it should be easier to get grants if you’re better at finding the truth. But in practice, unless someone does so badly they trigger a scandal, academics don’t usually get things all that wrong. So grants are mostly based on other factors.

Paying someone purely to deliver the truth, not to entertain or match a culture, seems tricky. You could imagine sci-fi scenarios. What if we could track the logic people used to make decisions, and demand payment if those decisions were based on facts we uncovered, like a journalist getting a percentage of every short made in response to bad news they dug up about a company? What if governments paid in proportion to how valuable academic ideas turned out to be, centuries after they were discovered, and modern-day academics sold shares in that future payout to fund themselves? What if prediction markets something something?

For the moment, academics and journalists are both in a weird middle space. They’re truthseekers, still, by culture and inclination and desire. But they’re paid for something else.

September 03, 2026

Tommaso Dorigo — When A Bump Gets Greedy: A Connection I Had Missed For 25 Years

When A Bump Gets Greedy: A Connection I Had Missed For 25 Years

Back in 2009 I wrote in this blog a rather technical post about something I had called the “greedy bump bias.” (GBB) The effect I referred to had emerged from a 2003 CDF study of mine, where I considered the extraction of small signals sitting on top of much larger backgrounds.That work went unpublished, but then I took revenge with the blog post...

Tommaso Dorigo
Categories

September 01, 2026

n-Category Café Three Generations in E7

It’s long been a mystery why there are 3 generations of quarks and leptons: three sets of particles, apparently identical except for how they interact with the Higgs boson. It would be nice if there were some good physical explanation. Nobody knows one. Barring that, it would be nice if some beautiful mathematical structure made this pattern seem natural. That’s what my new paper is about.

I’ll keep this nontechnical. I’ll say a bit about what the paper does, what it does not do, what led up to it, and how I wrote it.

This is my third paper about exceptional algebraic structures and the Standard Model. When you classify famous gadgets in algebra, beautiful gadgets with fancy names like ‘simple Lie algebras’ and ‘Euclidean Jordan algebras’ and ‘positive hermitian Jordan pairs’, you tend to get infinite series of them — together with a few exceptions that can be built using the octonions. This is a bit spooky, so I’ve been interested in this for a long time.

A few physicists have hoped that these exceptions are good for something. For example, maybe the quirky features of our best theory of particle physics, the Standard Model, aren’t accidental. Perhaps they fall out naturally from some exceptional algebraic structure.

It’s a long shot, but we’ve been stuck on figuring out new fundamental laws of particle physics for so long — roughly since the early 1980s — that it’s worth a try.

In 2018, Michel Dubois-Violette and Ivan Todorov noticed that the gauge group of the Standard Model falls out as symmetries of the so-called ‘exceptional Jordan algebra’ together with some ordinary Jordan algebras sitting inside it. I tried to clarify that here, with a huge amount of help from an excellent young mathematician:

It’s very nice, because the Jordan algebras in question arise naturally when you try to axiomatize the foundations of quantum physics. It would be so cool if something about quantum physics made the Standard Model seem mathematically natural!

But really this result only concerns the gauge bosons in the Standard Model: the photon, gluons, and the W and Z bosons. It says nothing about the fermions — that is, the quarks and leptons. And it seems quite hard to get those into the picture.

In 2020, Latham Boyle tried to solve this problem by tensoring the exceptional Jordan algebra with the complex numbers. This made one generation of fermions appear quite naturally! But the connection to the foundations of quantum physics seemed lost: tensoring the exceptional Jordan algebra with the complex numbers seems at first like it might be just a formal trick.

This spring, Latham and his student Endre Bokor and I showed the connection to quantum physics is not lost:

The idea is to work, not with Jordan algebras, but with more general things called Jordan pairs, which have been studied by mathematicians since at least 1975. We showed that you can still do quantum physics with Jordan pairs. And we showed that there’s an ‘exceptional’ Jordan pair that naturally contains the Standard Model gauge group and one generation of fermions!

This Jordan pair is built from the bioctonions: the octonions tensored with the complex numbers. And it’s closely related to an exceptional Lie algebra called 𝔢 6\mathfrak{e}_6.

This is nice because the work of Dubois-Violette and Todorov used a smaller exceptional Lie algebra called 𝔣 4\mathfrak{f}_4. Going up to 𝔢 6\mathfrak{e}_6 gives the room to include one generation of fermions.

There’s an even larger exceptional Lie algebra you can use to build a Jordan pair: it’s called 𝔢 7\mathfrak{e}_7. Bokor, Boyle and I tried using this to get three generations of fermions. There are things that make this tempting: not just the fact that 𝔢 7\mathfrak{e}_7 is bigger, but the fact that the Jordan pair you get from it has a kind of three-fold symmetry. But we couldn’t get it to work.

Around this time I got very interested in some work that someone had sent me in October 2025. My inbox is packed with new theories of physics. Since the rise of large language models the inflow has increased: I get about two emails a day from someone telling me they’ve made a revolutionary discovery in physics. Practically none of these theories appeal to me. But this paper, and this thesis, were different:

He claimed to fit three generations of fermions into the exceptional Lie algebra 𝔢 7\mathfrak{e}_7.

When I started seriously trying to understand this paper, I wound up translating it into a language I’m more comfortable with, and expanding on the ideas a bit. So I wrote this:

Here’s the basic idea.

The idea

There is a standard way to fit the Lie algebra of the Standard Model gauge group, which I call 𝔤 SM\mathfrak{g}_{\text{SM}}, into the Lie algebra 𝔢 7\mathfrak{e}_7. You can construct a Lie algebra LL that fits between them:

𝔤 SM⊂L⊂𝔢 7 \mathfrak{g}_{\text{SM}} \subset L \subset \mathfrak{e}_7

As a vector space we have

𝔢 7≅L⊕V \mathfrak{e}_7 \; \cong \; L \oplus V

for some vector space VV of dimension 3×323 \times 32.

Moreover, the Lie algebra 𝔤 SM\mathfrak{g}_{\text{SM}} acts on VV, via the 𝔢 7\mathfrak{e}_7 Lie bracket, precisely as it does on three generations of Standard Model fermions and their antiparticles, including right-handed neutrino and its antiparticle — but ignoring spin!

There is, in fact, a very interesting three-fold symmetry built into 𝔢 7\mathfrak{e}_7, which is revealed when we put the Standard Model Lie algebra 𝔤 SM\mathfrak{g}_{\text{SM}} into it. It permutes the three generations.

Like Nasmith, I am not proposing a theory of physics. I’m only observing a fascinating mathematical pattern that might (or might not) be of some use in physics.

There are lots of things this pattern does not include: basically, everything I didn’t already mention. It does not include the spin of the fermions and gauge bosons. It does not include the Higgs boson, though in some sense it comes close (see the paper). It does not include a Lagrangian, so it doesn’t say anything at all about particle masses or interactions.

I could say a lot more about what my paper does do… most importantly, where this Lie algebra LL comes from! The details are very interesting. There’s also the curious role of the right-handed neutrinos. But I’ve already spent weeks explaining all these things in my paper, so I won’t do it here. Instead let me say a bit about how I wrote the paper.

Writing the paper

I’ve been wanting to keep up with how AI is transforming math. About a year ago a friend gave me a subscription to Claude Pro. I wanted to test it out, despite my many misgivings, including how large language models are contributing to global warming and income inequality. Given the amazing things that people have recently done in math using large language models, I didn’t think that never trying them out would put me in the best position to make good decisions about the future.

So, I wrote this paper with help from Claude Opus 4.8.

I started by giving it Nasmith’s paper and asking a long series of questions about that paper over several days. The results were very interesting and helpful. Eventually I asked it to summarize and expand on our conversation. It quickly spat out a 10-page paper.

This paper was written in a breezy, pleasant style — but also quite hard to understand in detail, since it mixed Nasmith’s terminology with the Lie algebra terminology I prefer, and the proofs skipped over some steps.

It took me about three weeks of hard work to fully understand and re-express all the ideas a way that I like. For a while I felt dumb and frustrated, because when I asked Claude to fill in the gaps in proofs, it used math I was not very competent in, like the theory of regular subalgebras, and the theory of minuscule representations. But I learned this math, and everything turned out to be basically correct — in part, I’m sure, because Nasmith’s original work was correct.

For several weeks I checked, reorganized, expanded and completely rewrote this material. By the end everything was written in a style I like, emphasizing the ideas I consider important, proving things fairly carefully, and adding a lot of expository material — for example, explaining the theory of regular subalgebras.

Almost no traces of Claude’s original writeup remain, even though I was deeply influenced by them. My proofs make few references to deep theorems, though they assume solid familiarity with simple Lie algebras and their root systems. The proofs also require no brutally hard computations — though Claude was eager to do such computations to check things.

Any mistakes in this paper are my own.

I’m not sure what conclusions I draw from writing this paper. I’m writing another math paper now, with a human coauthor, and I have no desire to get help from a large language model. For work on my own it could be very helpful. Fields medalist Jacob Tsimerman says it roughly doubles his productivity. Would using it be so bad for the environment, or so bad for society, that I should avoid it? Maybe. I deliberately stuck with Claude Opus 4.8 instead of something more powerful, to see what I could do with what you get from a $20/month subscription. But maybe that’s still bad.

I avoid flying to conferences, which in some ways cripples my ability to keep up with new trends and influence people — but I don’t mind that. It gives me more time to think.

I will think carefully about my next move.

August 31, 2026

John Preskill — The Universe, the Uncanny, and Fashion

                                                                  ⚛⚛⚛

In the beginning there was a question. Actually, no, in the beginning there was language.  The question came later, presumably as one of its side effects. Since then we have asked about nearly everything, though the answers have done little to alter our circumstances. Human existence has always seemed strange to me. We arrive without consent, in a place not of our choosing, then spend decades asking why, until we close our eyes and enter the abyss of nothingness.

Some people accept this situation very well. I did not.

 Until my mid-teens, I tenaciously found one solution in fashion. If I could not understand who I was, I could at least decide what version I could become. Clothes gave form to something otherwise difficult to locate. Desire, after all, begins with a lack. 

It was only later that physics presented a solution to the same lack, only with more elaborate mathematics.

Quantum mechanics tells us that the world beneath the familiar world does not behave as we experience it.  A thing can resist being one thing, certainty begins to dissolve, and reality becomes strangely unfamiliar. The uncanny begins very close to home.

It is in this uncanniness that fashion and physics, liminally, meet for me: both begin with a human being standing before something they cannot understand, the universe in one case, the self in the other, and trying to make a form out of it. 

But even this distinction, on closer inspection, begins to collapse; for our desire to understand the universe has always concealed a deeper desire to understand the self that stands within it. 

All roads eventually lead back to the interests you had as a child, or so I say. 

I would squint into the dark, my mind already at work. I would imagine the next thing I would wear, project its color and fabric onto the ceiling. 

Then,  go to the bazaar, Rangrizano Dana, where, as the Kandaharis like to say, everything is sold except one’s mother and father.

The designs that had existed only in my head, after night upon night of theorizing and calculating, were finally beginning to take shape. The alleys of Rangrizano Dana were full of fabric of different textures and colors. 

Once the fabric was bought, came the tailoring. Tailors hold a special place in Afghanistan. Some are stars in their own right, the kind you have to book an appointment with. And once you get there, another matter comes up: making the tailor promise not to show your design to another woman. Everyone wants to wear something unique. So, naturally,  there is a lot of secrecy.

The designs themselves would either be described or sketched. Somewhere between what was imagined and what the tailor could make, something would emerge. And the day the tailor shipped your clothes was really a day of revelations.

Time passed, which is another way of saying the object of my desire changed. The same attention once spent on cut, texture, and appearance turned, little by little, toward the fabric of space and time.

Studying physics, I organized my process in the same way. There was again an idea that existed first in the mind, and then the problem of giving it form. 

Only now the materials were different. Instead of cloth, there was mathematics; instead of the tailor’s table, the blackboard and chalk. And, standing before it, more often than not, a badly dressed physicist covered in chalk dust. All the libidinal energy of the universe, you might say, had been sublimated into equations.

Fashion had, in some ways, been the more serious pursuit. There was something at stake: you wanted to look better than everyone else. Physics, by comparison, could be surprisingly playful.

Alice sends a particle to Bob. The sound of those names would tickle my Pashto ear then, and still cracks me up now. Cats are placed in boxes, both dead and alive. Observers hover near black holes or sometimes even, whooooooosh, fall into them.

So, in this way, in my mind’s eye, the physicist has become a rather more amusing figure: a tailor of space and time. The universe, of course, is a notoriously difficult client (perhaps even more difficult than a Kandahari woman making her tailor promise not to show her design to anyone else). 

He begins, as any tailor must, with an imagined shape. Mathematics is his chalk, his scissors, his needle. With it he marks the fabric, cuts it, folds it, joins one piece of reality to another. Some constructions fall upon the universe with elegance. Others bunch at the shoulder, pull at the waist, or refuse altogether to button, and must be altered, or thrown away completely.

A tailor works against the resistance of the human body; a physicist against the resistance of the universe.

Imagine, then, a physicist seated at the edge of a black hole, spectacles low upon his nose, sewing little cloaks from the fabric of spacetime.

“How many dimensions will you need?” asks Alice.

“That depends upon the suit,” he replies.

    For AdS/CFT, suppose we are making a three-dimensional garment.  The tailor of the universe never begins with three dimensions. Before him lies only the flat, two-dimensional fabric, with all its degrees of freedom spread upon the board.

 It is then, in a less sartorial language, we might call the disentangling and entangling of degrees of freedom; the tailor coaxes a third dimension out of the flat cloth. This is the secret embroidery of the holographic idea: what appears to the wearer as a three-dimensional world is encoded in a fundamentally two-dimensional way.

Bob tries this newly fashioned quantum garment on before entering the black hole. 

 “Too tight,” he says. “See if there is more degree of freedom.” 

The physicist frowns, takes up his chalk, and makes another small mark, with some disentangling and entangling here and there.

What could have been wrong? Well, perhaps the group theory has been chosen badly, the symmetry is broken, and a representation must be changed, a seam opened, a dimension added, or one cunningly concealed.

Alice watches him work, the almighty tailor of the universe,  making change without a change.

“But how do you know when the suit is right?”

And with that, the physicist looks at her.

“What do you want? Do you want to put me out of work? Unknowing is the lack that drives us to make more suits. If we knew the suit was right, that would be the end of everything.”

August 28, 2026

Matt von Hippel — Don’t Judge an Explanation by Its Cover

Dark matter bugs people.

I’ve talked before about why, and why it, and other beyond-the-standard-model proposals like those inspired by MOND, are nonetheless credible with physicists. But beyond the logic in that post, there’s a deeper reason people find dark matter strange. It’s that they don’t know what kind of an explanation dark matter is.

Dark matter sounds very lazy. If you can’t explain the movements of stars based on the matter you can see, then proposing invisible matter sounds like the easy way out. But it’s actually a lot less easy than it sounds, because matter is something quite specific. Matter gravitates and bends light. Matter moves. Matter can be described with a pressure, one like gas and dust and not like other things like light or the Higgs field. If you propose a new type of matter, you have to check and see that all of those consequences hold, with detailed implications for almost every observation every astronomer takes.

For the most part, those consequences have been checked, and they do hold. Sometimes they fail, and it’s those failures, and not the idea that dark matter is “lazy”, that drive dark matter’s critics in the physics profession. Physicists who oppose dark matter have other explanations with their own consequences, for example new types of quantum fields that often get described to the public as “modified gravity”. When they argue against dark matter, they do it by comparing those consequences in detail, working through the implications and seeing which phenomena hold.

Dark matter, as it turns out, is a very constraining explanation, one with strict consequences. There are other corners of physics where the explanations may seem less lazy, but actually have fewer consequences, and thereby less scientific heft.

For example, consider the debate about evidence for dark energy I wrote about last month. A key question there was how to interpret light from supernovae. Some groups argued that supernovae change in brightness with distance, others that they change based on how old their galaxies are. Sabine Hossenfelder glossed the debate by saying it comes down to how you model supernovae. And while that’s true, it can give the wrong impression.

You might think that these people are comparing detailed computer models of supernovae, and making different assumptions when they set their models up. But in reality, it’s much less detailed. The people on both sides of this debate are looking at correlations, trying to draw statistical lines through supernova datasets. The difference between one model and another isn’t a complicated physical setup you can put into a simulation, it’s just which lines on a graph you account for and which you ignore.

Because of that, while these models may sound much more sophisticated than dark matter, they actually have much less scientific weight. The different supernova models don’t have grand, widespread consequences, they’re not mucking with the laws of physics or proposing new classes of object that every astronomer needs to account for. They’re pretty much just proposing tweaks to how to interpret one very specific type of data. That makes their questions much harder to resolve, and their answers much less universally convincing.

If you’re not a scientist, if you read science news, it can be hard to tell the difference. Some ideas in science may sound simple, but have a whole raft of consequences that distinguish them from other ideas. Others may sound sophisticated, but are much more like “fudge factors”, only distinguished by statistical arguments, not by a rich trail of qualitative evidence.

For the most part, as an outsider, you’ll never know which is which. But as always, it’s best to be aware of your limits.

August 17, 2026

John Baez — Three Generations in E7

It’s long been a mystery why there are 3 generations of quarks and leptons: three sets of particles, apparently identical except for how they interact with the Higgs boson. It would be nice if there were some good physical explanation. Nobody knows one. Barring that, it would be nice if some beautiful mathematical structure made this pattern seem natural. That’s what my new paper is about.

It’s my third paper about exceptional algebraic structures and the Standard Model. When you classify famous gadgets in algebra, beautiful gadgets with fancy names like ‘simple Lie algebras’ and ‘Euclidean Jordan algebras’ and ‘positive hermitian Jordan pairs’, you tend to get infinite series of them—together with a few exceptions that can be built using the octonions. This is a bit spooky, so I’ve been interested in this for a long time.

A few physicists have hoped that these exceptions are good for something. For example, maybe the quirky features of our best theory of particle physics, the Standard Model, aren’t accidental. Perhaps they fall out naturally from some exceptional algebraic structure.

It’s a long shot, but we’ve been stuck on figuring out new fundamental laws of particle physics for so long—roughly since the early 1980s—that it’s worth a try.

In 2018, Michel Dubois-Violette and Ivan Todorov noticed that the gauge group of the Standard Model falls out as symmetries of the so-called ‘exceptional Jordan algebra’ together with some ordinary Jordan algebras sitting inside it. I tried to clarify that here, with a huge amount of help from an excellent young mathematician:

• John Baez and Paul Schwahn, The Standard Model gauge group from the exceptional Jordan algebra. (Blog article here.)

It’s very nice, because the Jordan algebras in question arise naturally when you try to axiomatize the foundations of quantum physics. It would be so cool if something about quantum physics made the Standard Model seem mathematically natural!

But really this result only concerns the gauge bosons in the Standard Model: the photon, gluons, and the W and Z bosons. It says nothing about the fermions—that is, the quarks and leptons. And it seems quite hard to get those into the picture.

In 2020, Latham Boyle tried to solve this problem by tensoring the exceptional Jordan algebra with the complex numbers. This made one generation of fermions appear quite naturally! But the connection to the foundations of quantum physics seemed lost: tensoring the exceptional Jordan algebra with the complex numbers seems at first like it might be just a formal trick.

This spring, Latham and his student Endre Bokor and I showed the connection to quantum physics is not lost:

• John Baez, Endre Bokor and Latham Boyle, Jordan pair quantum theory and the Standard Model. (Blog article here.)

The idea is to work, not with Jordan algebras, but with more general things called Jordan pairs, which have been studied by mathematicians since at least 1975. We showed that you can still do quantum physics with Jordan pairs. And we showed that there’s an ‘exceptional’ Jordan pair that naturally contains the Standard Model gauge group and one generation of fermions!

This Jordan pair is built from the bioctonions: the octonions tensored with the complex numbers. And it’s closely related to an exceptional Lie algebra called \mathfrak{e}_6.

This is nice because the work of Dubois-Violette and Todorov used a smaller exceptional Lie algebra called \mathfrak{f}_4. Going up to \mathfrak{e}_6 gives the room to include one generation of fermions.

There’s an even larger exceptional Lie algebra you can use to build a Jordan pair: it’s called \mathfrak{e}_7. Bokor, Boyle and I tried using this to get three generations of fermions. There are things that make this tempting: not just the fact that \mathfrak{e}_7 is bigger, but the fact that the Jordan pair you get from it has a kind of three-fold symmetry. But we couldn’t get it to work.

Around this time I got very interested in some work that someone had sent me in October 2025. My inbox is packed with new theories of physics. Since the rise of large language models the inflow has increased: I get about two emails a day from someone telling me they’ve made a revolutionary discovery in physics. Practically none of these theories appeal to me. But this paper, and this thesis, were different:

• Benjamin Nasmith, An exceptional combinatorial sequence and Standard Model particles, 2020.

• Benjamin Nasmith, Tight Projective 5-Designs and Exceptional Structures, Ph.D. thesis, Royal Military College of Canada, 2023.

He claimed to fit three generations of fermions into the exceptional Lie algebra \mathfrak{e}_7.

When I started seriously trying to understand this paper, I wound up translating it into a language I’m more comfortable with, and expanding on the ideas a bit. So I wrote this:

• John Baez, Three generations in \mathfrak{e}_7.

Here’s the basic idea.

The idea

There is a standard way to fit the Lie algebra of the Standard Model gauge group, which I call \mathfrak{g}_{\text{SM}}, into the Lie algebra \mathfrak{e}_7. You can construct a Lie algebra L that fits between them:

\mathfrak{g}_{\text{SM}} \subset L  \subset \mathfrak{e}_7

As a vector space we have

\mathfrak{e}_7 \; \cong \; L \oplus V

for some vector space V of dimension 3 \times 32.

Moreover, the Lie algebra \mathfrak{g}_{\text{SM}} acts on V, via the \mathfrak{e}_7 Lie bracket, precisely as it does on three generations of Standard Model fermions and their antiparticles, including right-handed neutrino and its antiparticle—but ignoring spin!

There is, in fact, a very interesting three-fold symmetry built into \mathfrak{e}_7, which is revealed when we put the Standard Model Lie algebra \mathfrak{g}_{\text{SM}} into it. It permutes the three generations.

Like Nasmith, I am not proposing a theory of physics. I’m only observing a fascinating mathematical pattern that might (or might not) be of some use in physics.

There are lots of things this pattern does not include: basically, everything I didn’t already mention. It does not include the spin of the fermions and gauge bosons. It does not include the Higgs boson, though in some sense it comes close (see the paper). It does not include a Lagrangian, so it doesn’t say anything at all about particle masses or interactions.

I could say a lot more about this… most importantly, where this Lie algebra L comes from. The details are very interesting. There’s also the curious role of the right-handed neutrinos. But I’ve already spent weeks explaining all these things in my paper, so I won’t do it here. Instead let me say a bit about how I wrote the paper.

Writing the paper

I’ve been wanting to keep up with how AI is transforming math. About a year ago a friend gave me a subscription to Claude Pro. I wanted to test it out, despite my many misgivings, including how large language models are contributing to global warming and income inequality. Given the amazing things that people have recently done in math using large language models, I didn’t think that never trying them out would put me in the best position to make good decisions about the future.

So, I wrote this paper with help from Claude Opus 4.8.

I started by giving it Nasmith’s paper and asking a long series of questions about that paper over several days. The results were very interesting and helpful. Eventually I asked it to summarize and expand on our conversation. It quickly spat out a 10-page paper.

This paper was written in a breezy, pleasant style—but also quite hard to understand in detail, since it mixed Nasmith’s terminology with the Lie algebra terminology I prefer, and the proofs skipped over some steps.

It took me about three weeks of hard work to fully understand and re-express all the ideas a way that I like. For a while I felt dumb and frustrated, because when I asked Claude to fill in the gaps in proofs, it used math I was not very competent in, like the theory of regular subalgebras, and the theory of minuscule representations. But I learned this math, and everything turned out to be basically correct—in part, I’m sure, because Nasmith’s original work was correct.

For several weeks I checked, reorganized, expanded and completely rewrote this material. By the end everything was written in a style I like, emphasizing the ideas I consider important, proving things fairly carefully, and adding a lot of expository material—for example, explaining the theory of regular subalgebras.

Almost no traces of Claude’s original writeup remain, even though I was deeply influenced by them. My proofs make few references to deep theorems, though they assume solid familiarity with simple Lie algebras and their root systems. The proofs also require no brutally hard computations—though Claude was eager to do such computations to check things.

Any mistakes in this paper are my own.

I’m not sure what conclusions I draw from writing this paper. I’m writing another math paper now, with a human coauthor, and I have no desire to get help from a large language model. For work on my own it could be very helpful. Jacob Tsimerman says it roughly doubles his productivity. Would using it be so bad for the environment, or so bad for society, that I should avoid it? Maybe. I deliberately stuck with Claude Opus 4.8 instead of something more powerful, to see what I could do with what you get from a $20/month subscription. But maybe that’s still bad.

I avoid flying to conferences, which in some ways cripples my ability to keep up with new trends and influence people—but I don’t mind that. It gives me more time to think.

I will think carefully about my next move.

John Baez — Jordan Triples and the Standard Model

I don’t usually talk about particle physics here. I have a whole series of articles about octonions and the Standard Model on my other blog. But I’m kind of excited about this new paper, so I’ll talk about it here too:

• John Baez, Endre Bokor and Latham Boyle, Jordan pair quantum theory and the Standard Model.

Jordan algebras were introduced by Jordan, von Neumann and Wigner in 1934 in an attempt to formalize algebras of observables in quantum theory. They come in 4 infinite series—but there’s one more, the ‘exceptional Jordan algebra’, consisting of 3 × 3 self-adjoint matrices of octonions. For years physicists sought to find some use for it.

In 2018, Todorov and Dubois–Violette noticed that the symmetries of the exceptional Jordan include the Standard Model gauge group in a nice way. But it was unclear how to bring in the fermions—the quarks and leptons. That’s what our new paper does.

To do this, we need to go beyond Jordan algebras. Jordan pairs and Jordan triples are two closely linked formalisms that generalize Jordan algebras. Our paper explains them in detail—and how they’re connected to geometry and quantum mechanics. But here I will mostly skip that wonderful story, so I can quickly explain the connection to the Standard Model.

Here’s how the Standard Model gauge group, together with its representation on one generation of fermions, drops out of a Jordan triple.

The bi-Cayley triple

Let

\mathbb{O}_\mathbb{C} = \mathbb{C} \textstyle{\otimes}_\mathbb{R} \mathbb{O}

be the bioctonions: octonions with complex coefficients. Write \mathbb{O}_\mathbb{C}^2 for the space of column vectors with two bioctonion entries.

\mathbb{O}_\mathbb{C}^2 has a certain triple product

[x,y,z]=\frac{1}{2}(x(y^{\dagger}z)+z(y^{\dagger}x))

which obey the axioms of a gadget called a ‘positive hermitian Jordan triple’. It’s called the bi-Cayley triple.

Now, every positive hermitian Jordan triple gives rise to a \mathbb{Z}_2-graded real Lie algebra

\mathbf{k} = \mathbf{k}_0 \textstyle{\oplus} \mathbf{k}_1

Not a Lie superalgebra: a plain old-fashioned Lie algebra with a \mathbb{Z}_2-grading!

How does this work? We take the hermitian Jordan triple itself to be \mathbf{k}_1. The Lie algebra \mathbf{k}_0 consists of all linear maps from \mathbf{k}_1 to itself that are of this form:

x \mapsto [a,b,x] - [b,a,x]

for some a,b \in \mathbf{k}_1. These maps are called real inner derivations. They form a Lie algebra since the commutator of two such maps is another such map. With a bit more work we can define other operations making all of \mathbf{k} into a \mathbb{Z}_2-graded Lie algebra.

So, we get a big Lie algebra \mathbf{k}, and a Lie subalgebra \mathbf{k}_0 sitting inside it. From this we get two Lie groups: a big one K whose Lie algebra is \mathbf{k}, and a subgroup K_0 whose Lie algebra is \mathbf{k}_0.

The quotient is K/K_0 is a nice kind of manifold called a hermitian symmetric space. Conversely, any compact hermitian symmetric space give rise to a positive hermitian Jordan triple!

This geometric picture is revealing. The group K acts transitively as symmetries of our hermitian symmetric space, while the stabilizer of any point is isomorphic to K_0. Our original Jordan triple, \mathbf{k}_1, is then the tangent space of that point! So, K_0 acts on this Jordan triple. This action preserves the triple product, and we call K_0 the real inner automorphism group of our Jordan triple.

Here’s another great thing about the geometric picture: hermitian symmetric spaces were classified by Eli Cartan (who seems to have spent his life classifying things). As a result we also know the classification of positive hermitian Jordan triples. They come in four infinite series together with two exceptions. One is the bi-Cayley triple, and other is the Albert triple, which is the complexification of the exceptional Jordan algebra. The bi-Cayley triple is a subtriple of the Albert triple. It’s these two exceptions that are connected to the Standard Model. But we’ll start with the bi-Cayley triple.

The 3-graded Lie algebra coming from the bi-Cayley triple is the compact real form of \mathfrak{e}_6:

\mathfrak{e}_6 = \big[\mathfrak{so}(10) \textstyle{\oplus} \mathfrak{u}(1)\big] \textstyle{\oplus} \mathbb{O}_\mathbb{C}^2

The even part of this Lie algebra is in brackets. The corresponding hermitian symmetric space is called the bioctonionic plane (\mathbb{C}\otimes\mathbb{O})P^2. The even part of our 3-graded Lie algebra, \mathfrak{so}(10)\oplus \mathfrak{u}(1), generates the stabilizer of a point in the bioctonionic plane. The odd part, our friend \mathbb{O}_\mathbb{C}^2, is the tangent space of that point.

Here’s the first big surprise. The even part transforms as the adjoint representation of \mathrm{Spin}(10), while the odd part itself transforms as the 16-dimensional complex spinor representation of \mathrm{Spin}(10). Ignoring the extra \mathrm{U}(1) for a moment, this is exactly what we see in a \mathrm{SO}(10) grand unified theory: gauge bosons in the adjoint representation, and one generation of fermions in the 16-dimensional spinor representation.

So before we do anything, the bi-Cayley triple already smells like it contains the ingredients of an \mathrm{SO}(10) grand unified theory.

Tripotents

In a Jordan algebra the important elements are the idempotents, e^2 = e. In a Jordan triple W their role is played by tripotents: elements e with

[e,e,e] = e

A tripotent always lets us split W into three parts via something called its Peirce decomposition. The operator w \mapsto [e,e,w] has eigenvalues 0, 1/2, and 1, so W splits into the corresponding eigenspaces

W = W_0(e) \textstyle{\oplus} W_{1/2}(e) \textstyle{\oplus} W_1(e)

which are called the Peirce 0-space, Peirce 1/2-space and Peirce 1-space of e. A tripotent is called minimal when its Peirce 1-space is one-dimensional. Two tripotents e_1, e_2 are called colinear when each lies in the other’s Peirce 1/2-space.

I can’t resist explaining some of the quantum physics here. In a hermitian Jordan triple, the triple product [-,-,-] is linear in the first and last slot, but conjugate-linear in the middle slot. So, if you multiply a tripotent by a phase \alpha, you get a new tripotent:

[\alpha e, \alpha e, \alpha e] = \alpha \overline{\alpha} \alpha e = \alpha e

This should remind you of how when you multiply a unit vector in a Hilbert space by a phase, you get a new unit vector. In Jordan triple quantum mechanics, minimal tripotents take the place of these unit vectors. The hermitian symmetric space K/K_0 that I was talking about earlier is the same as the space of minimal tripotents mod phase! So, it generalizes the familiar space of ‘pure states’ in quantum mechanics: unit vectors mod phase.

But let’s get back to the Standard Model.

A chain of Jordan triples

From here on, the single fact driving everything is this: in any hermitian Jordan triple, any minimal tripotent’s Peirce 1/2-space is itself a hermitian Jordan triple!

If we run this starting from the bi-Cayley triple, we get this chain of hermitian Jordan triples, where each row’s 1/2-space is the next row’s triple:

Jordan triple Lie algebra \mathbf{k}_0 \oplus \mathbf{k}_1 (even part in brackets)
W = \mathbb{O}_\mathbb{C}^2 \mathfrak{e}_6 = [\mathfrak{so}(10) \oplus \mathfrak{u}(1)] \oplus \mathbb{O}_\mathbb{C}^2
W' = \mathfrak{a}_5(\mathbb{C}) \mathfrak{so}(10) = [\mathfrak{su}(5) \oplus \mathfrak{u}(1)] \oplus \mathfrak{a}_5(\mathbb{C})
W'' = \mathrm{M}_{3,2}(\mathbb{C}) \mathfrak{su}(5) = [\mathfrak{g}_{\mathrm{SM}}] \oplus \mathrm{M}_{3,2}(\mathbb{C})

Here \mathfrak{a}_5(\mathbb{C}) is the Jordan triple of antisymmetric 5\times 5 complex matrices, \mathrm{M}_{3,2}(\mathbb{C}) is the Jordan triple of 3\times 2 complex matrices, \mathfrak{g}_{\mathrm{SM}} = \mathfrak{su}(3)\oplus\mathfrak{su}(2)\oplus \mathfrak{u}(1), and

G_{\mathrm{SM}} = \mathrm{S}(\mathrm{U}(2) \times \mathrm{U}(3)) \cong (\mathrm{SU}(3)\times\mathrm{SU}(2)\times\mathrm{U}(1))/\mathbb{Z}_6

is the true Standard Model gauge group.

The gauge group from two tripotents

Start with the bi-Cayley triple. Choose two colinear minimal tripotents e_1, e_2. Descend the table twice:

• Start with W = \mathbb{O}_\mathbb{C}^2, which has real inner automorphism group (\mathrm{Spin}(10)\times\mathrm{U}(1))/\mathbb{Z}_4.

• Fix e_1. Its Peirce 1/2-space is W' = \mathfrak{a}_5(\mathbb{C}), with real inner automorphism group \mathrm{SU}(5)\times\mathrm{U}(1).

• Fix e_2 (colinear with e_1, so living in W'). Its Peirce 1/2-space in W' is W'' = \mathrm{M}_{3,2}(\mathbb{C}), with real inner automorphism group exactly G_{\mathrm{SM}}.

In other words, the subspace of the bi-Cayley triple colinear with both e_1 and e_2 is a Jordan triple whose real inner automorphism group is the Standard Model gauge group.

The choice of e_1 and e_2 also pins down how G_{\mathrm{SM}} sits inside the original group \mathrm{E}_6. At each we step take the subgroup that acts with determinant 1 and preserves the chosen tripotent up to a phase; this gives a chain of subgroups whose members are \mathrm{Spin}(10), \mathrm{U}(5), and G_{\mathrm{SM}}, so we get the embeddings

G_{\mathrm{SM}} \subset \mathrm{SU}(5) \subset \mathrm{Spin}(10)

In particle physics, this is the classic chain taking us from the so-called \mathrm{SO}(10) grand unified theory down to the \mathrm{SU}(5) grand unified theory down to the Standard Model. And it’s well known that restricting the 16-dimensional complex spinor representation of \mathrm{Spin}(10) along this chain gives precisely the Standard Model representation \rho_{\mathrm{SM}} on one generation of fermions! So we get one generation of Standard Model fermions this way.

The six particles types as Peirce spaces

We have gotten the representation of the Standard Model gauge group on one generation of fermions without any fuss. But it’s also fun to peer into the details, and see how the different kinds of fermions emerge.

For any tripotent e, we have projections P_0(e), P_{1/2}(e) and P_1(e) onto its three eigenspaces: its so-called Peirce projectors. Since we get the Standard Model structure using two minimal tripotents e_1 and e_2 in the bi-Cayley triple \mathbb{O}_{\mathbb{C}}^2, there are nine composites of two Peirce projectors we can apply to this triple. This is how we pick out the different kinds of fermions!

As a representation of the Standard Model Lie algebra

\mathfrak{g}_{\mathrm{SM}} = \mathfrak{su}(3) \textstyle{\oplus} \mathfrak{su}(2) \textstyle{\oplus} \mathfrak{u}(1)

any generation of Standard Model fermions transforms as the direct sum of six irreducible representations:

\rho_{\mathrm{SM}} = (3,2,\tfrac{1}{6}) \textstyle{\oplus} (\bar 3,1,\tfrac{1}{3}) \textstyle{\oplus} (\bar 3,1,-\tfrac{2}{3}) \textstyle{\oplus} (1,2,-\tfrac{1}{2}) \textstyle{\oplus} (1,1,1) \textstyle{\oplus} (1,1,0)

These correspond to the six types of left-handed fermion: q_L, \overline{d_R}, \overline{u_R}, \ell_L, \overline{e_R}, \overline{\nu_R}. Six irreducible pieces, six particle types.

It turns out these are exactly the six nonzero components of the Peirce decomposition of \mathbb{O}_\mathbb{C}^2 with respect to both e_1 and e_2. Those six match up one-to-one with the particle types:

Peirce projector representation of G_{\text{SM}} particle type
P_{1/2}(e_2) P_{1/2}(e_1) (3, 2, +1/6) q_L
P_{1/2}(e_2) P_0(e_1) (\overline{3}, 1, +1/3) \overline{d_R}
P_0(e_2) P_{1/2}(e_1) (\overline{3}, 1, −2/3) \overline{u_R}
P_0(e_2) P_0(e_1) (1, 2, −1/2) \ell_L
P_1(e_2) P_{1/2}(e_1) (1, 1, +1) \overline{e_R}
P_{1/2}(e_2) P_1(e_1) (1, 1, 0) \overline{\nu_R}

The remaining three combinations—P_1(e_2)P_1(e_1), P_1(e_2)P_0(e_1), and P_0(e_2)P_1(e_1)—all vanish, which is why we land on six pieces and not nine.

So the whole package—the gauge group G_{\mathrm{SM}}, the embedding G_{\mathrm{SM}} \subset \mathrm{Spin}(10), the representation \rho_{\mathrm{SM}}, and even the split of one generation into its six particle multiplets as distinct Peirce components—all comes out of the single object \mathbb{O}_\mathbb{C}^2 once you choose two colinear minimal tripotents.

And if you prefer to start one level up, with the Albert triple \mathfrak{h}_3(\mathbb{O}) \otimes \mathbb{C}, you get the same result by choosing three mutually colinear tripotents instead of two—but for that, read our paper!

August 13, 2026

Tim Gowers — What sort of maths are LLMs good at?

For the sake of anyone who might read this blog post in the distant future (a month from now, say), let me mention that I am writing it a few days after OpenAI announced that it had solved ten major problems in mathematics and theoretical computer science, including the first construction of a non-sofic group, and a proof that the multicolour Ramsey number R(3,3,...,3) (where there are k 3’s) grows superexponentially in k. The first was, to judge from various talks I have been to, one of the most important unsolved problems in group theory, and the second was a major open problem in Ramsey theory that I didn’t necessarily expect to see solved in my lifetime, though of course such expectations now have to be revised. The reason I want to be clear about the timing is that I shall be discussing the current capabilities of LLMs in the full expectation that those will continue to change rapidly. So it is likely that in not too long from now, if there is anything interesting in what I write, it will be interesting mainly as a record of what the situation looked like in early August 2026.

These results, and the other eight on the list, are extraordinarily impressive, but it still doesn’t seem to be the case that LLMs are better than all humans at all aspects of mathematics. If they were, then their big speed advantage over us would mean that there would be much more of a flood of results. So it is natural to wonder about what kinds of problems LLMs are good at, and about where there is still room for improvement. I don’t pretend to have a good answer to this question, where a good answer would be a crisp classification that would fit the current examples well, but it is an interesting exercise to try to rule out some bad answers, and to try to identify potential answers that aren’t obviously contradicted by the evidence.

Are LLMs particularly good at finding counterexamples?

A first remark here is that LLMs are not just good at finding counterexamples: they can find proofs of difficult statements as well. However, it is notable that the most famous problems they have solved have almost all been with counterexamples rather than proofs. That is true of the two problems mentioned above, and also of the Jacobian conjecture and the unit distance conjecture.

If one wants to theorize that LLMs are particularly good at finding counterexamples, then there are two things it would be good to do to make the theory more convincing. The first may sound unproblematic: it is to decide when solving a problem counts as finding a counterexample. Once that is sorted out, the second is to come up with a potential explanation of why LLMs would be particularly well suited to solving problems of that particular kind.

What does it mean to find a counterexample?

Why am I suggesting that it is not completely obvious what it means to find a counterexample? Surely, one might suggest, all it means is that you have a statement of the form “Every object of such and such a type has such and such a property,” and you exhibit an object of the given type that does not have the given property.

However, this doesn’t always work. Consider a famous result of Vinogradov, which states that every sufficiently large positive integer is a sum of three primes. The negation of this statement is (or is equivalent to) the statement that for every positive integer N there exists an integer n\geq N such that n is not a sum of three primes. In other words, it states that every positive integer N has a certain property. Seen in this light, Vinogradov found an example of a positive integer N that does not have the given property. Do we want to say that Vinogradov found a counterexample? Clearly not — the result should obviously be classified as a theorem and not a counterexample.

Thus, we cannot just naively say that LLMs are particularly good at negating universally quantified statements: there has to be something about the nature of the universal quantification. With the three-primes example, it is clear that Vinogradov did not think, “How am I going to find N with this property?” Rather, what he thought would have been more like, “I’ve got an integer n that is very large. How am I going to show that it is a sum of three primes?” In other words, all his focus would have been on the universally quantified n, with the existentially quantified N being a sort of afterthought once the details of the proof have been worked out.

In general, many interesting results, when they are stated formally, begin with an alternation of two or three (or more) quantifiers. The question then becomes to determine which is the first “interesting” quantified variable in some sense. Here’s another example to illustrate the point, from the theory of finite-dimensional normed spaces. I’ll give a few mathematical details for those curious, but if you don’t care about those, then you can skip the next three paragraphs and should get the gist of what I am saying about this example.

Let X and Y be two n-dimensional normed spaces and let T be a linear map from X to Y. We say that T is a C–isomorphism if there exists \lambda>0 such that \lambda\|x\|\leq\|Tx\|\leq C\lambda\|x\| for every x\in X. By rescaling we can always take \lambda to be 1, in which case we have that \|x\|\leq\|Tx\|\leq C\|x\| for every x\in X. If C=1, then this tells us that T is an isometry. In general, the Banach-Mazur distance d(X,Y) between X and Y is defined to be the smallest C such that there exists a C-isomorphism from X to Y. It is easy to see that the logarithm of the Banach-Mazur distance is a metric on the set of isometry classes of n-dimensional normed spaces. A less easy fact, but still not too hard, is that the resulting metric space is compact: in fact, it is known as the Banach-Mazur compactum.

It is natural to wonder what the diameter of the Banach-Mazur compactum is, and here things get interesting. A result of Fritz John states that every n-dimensional space X has distance at most \sqrt n from \ell_2^n. (The idea of the proof is as follows: pick inside the unit ball of X an n-dimensional ellipsoid of maximal volume; that is the unit ball of a normed space Y that is isometric to \ell_2^n; it can be shown that the identity map is a \sqrt n-isomorphism between X and Y.) From Fritz John’s theorem and the (multiplicative) triangle inequality, it follows that d(X,Y)\leq n for any two n-dimensional normed spaces. That is, the diameter of the Banach-Mazur compactum is at most n. But might it be substantially less than that?

An indication that the answer is not obvious comes from looking at the spaces \ell_1^n and \ell_\infty^n. The identity map between these two spaces is an n-isomorphism, but one can do much better by mapping the standard basis vectors not to themselves but to vertices of the unit cube, with the vertices chosen to be as orthogonal as possible. In particular, if there exists an n\times n Hadamard matrix, then the corresponding linear map is a \sqrt n-isomorphism. One can push this observation and deduce that for any p,q\in[1,\infty] the Banach-Mazur distance between \ell_p^n and \ell_q^n is O(\sqrt n). It is also easy to show that d(\ell_1^n,\ell_2^n)=\sqrt n, so \ell_p-spaces hardly improve on the easy lower bound, and do not improve on it at all in dimensions n for which an n\times n Hadamard matrix exists.

In 1981, Gluskin famously solved the problem by determining the correct asymptotics for the diameter of the Banach-Mazur compactum. Informally, what he showed was that the diameter is within a constant of the upper bound that follows immediately from Fritz John’s theorem. If we make the quantification explicit, then the statement we end up with is

\exists c>0\ \forall n\ \exists X,Y\in K_n\ d(X,Y)\geq cn,

where I have written K_n for the set of all n-dimensional normed spaces. (If you want to argue that it is not a set, then let me specify in addition that the underlying vector space is \mathbb R^n.) In words, there is a positive constant c such that for every positive integer n there are n-dimensional normed spaces X and Y such that the Banach-Mazur distance between X and Y is at least cn.

I can’t continue without very briefly describing the beautiful and highly influential idea Gluskin had for solving this problem. He took X and Y to be normed spaces whose unit balls were random symmetric convex sets defined as follows: take the standard basis vectors and a handful of other random unit vectors, as well as the negatives of all these vectors, and take the convex hull. Gluskin then showed that if two normed spaces are chosen from this distribution, then with high probability their Banach-Mazur distance is at least cn.

But back to the main point, which is that the logical form of the above statement is very similar to the logical form of Vinogradov’s theorem, which is

\exists N\ \forall n\geq N\ \exists p_1,p_2,p_3\in P\ \ p_1+p_2+p_3=n

where I have written P for the set of primes. And yet, Vinogradov’s result is unquestionably a theorem, while Gluskin’s result is unquestionably a counterexample, or at least an example.

What is the important difference between the two statements? It seems to be that in Vinogradov’s three-primes theorem the number n plays a more essential role in the statement that is to be proved about the various quantified variables. In Vinogradov’s theorem, that statement is n=p_1+p_2+p_3, whereas for Gluskin’s theorem the statement to be proved is

\dim X = \dim Y = n and d(X,Y)\geq cn,

which we can write equivalently as

\dim X = \dim Y = n and d(X,Y)\geq c\dim X.

In the case of Vinogradov’s theorem, the whole challenge is to get those three primes to add up to n, whereas for Gluskin it is not remotely challenging to get the dimensions of X and Y to equal n: the challenge is to get X and Y to be very far from each other, relative to their common dimension.

There is a further complication to bear in mind here, which is that via the process known as Skolemization, a universally quantified statement of the form \forall x\in X\ \exists y\in Y\ \ P(x,y) can be converted into an existentially quantifed statement \exists f:X\to Y\ \forall x\in X\ \ P(x,f(x)). (For this to be an equivalence one needs the axiom of choice, but it is certainly a sufficient condition.) This is not just a piece of logical trickery, but it often reflects quite accurately how we think about some problems. For instance, it is more natural to think of Gluskin’s example as a recipe for constructing (or at least proving the existence of) a pair of suitable normed spaces for any given dimension n, or in other words to construct a suitable function from \mathbb N to pairs of normed spaces by giving its value at each n, than it is to think of it as a statement that says that every positive integer n has a certain complicated property.

Yet another complication is that some universally quantified statements follow naturally from existentially quantified statements, or may even be equivalent to them. For example, the theorem that a 2-dimensional torus is not homeomorphic to a 2-dimensional sphere is a universally quantified statement (every map from the torus to the sphere fails to be a homeomorphism), but the natural way to prove it is to prove the existential statement that there is an invariant that distinguishes the two spaces. For an example of where a universal statement is equivalent to an existential statement, consider a statement of the form that a vector x\in\mathbb R^n does not belong to the convex hull of a certain compact set A. The statement that no convex combination of elements of A is equal to x is equivalent to the existence of a linear functional \phi:\mathbb R^n\to\mathbb R and a \lambda\in\mathbb R such that \phi(x)>\lambda and \phi(a)\leq\lambda for every a\in A. In both these cases it feels natural to regard the result as a theorem that is proved via an existential statement, perhaps because it is the theorem that is ultimately what interests us. But using “what interests us” as a criterion to determine what counts as a counterexample seems a little vague, and is a difficult criterion to use if we want to explain convincingly why AI should be good at finding counterexamples.

A more general argument against the notion that there is something about existential statements that is particularly suited to AI is that the need to establish existential statements pervades almost all of mathematical research, regardless of the nature of the headline result being aimed for. For example, if I want to prove a statement by induction, I may well look for a strengthening of the statement that serves better as an inductive hypothesis. Or if I want to prove that every object of type T with property P also has property Q, then I may well look for a property R that follows from P and can be used to prove Q. These are more metamathematical existence problems, but the distinction can be somewhat blurred, and more importantly, when trying to prove a statement S, it is often the case that the main question in our minds is less, “Why is S true?” and more, “What could a proof of S be like?” To give an example, I feel I understand pretty well why Goldbach’s conjecture is true — a highly plausible probabilistic model of the primes implies it and agrees closely with computational data — but if I were making a serious attempt to prove it, that understanding, which many mathematicians have had for a century or so, would be of limited help. Rather, my main task would be to try to find proof techniques that were powerful enough to make those heuristic ideas rigorous.

What is the difference between an example and a counterexample?

Logically, every statement of the form \exists x\ P(x) is a counterexample to the universally quantified statement \forall x\ \neg P(x). However, we do not describe all existential statements as counterexamples. For example, if I were to say, “The \ell_p-spaces with 1\leq p<\infty are all separable, as is c_0, but \ell_\infty is not separable,” I would not describe the second part of that assertion as a counterexample to the claim that all Banach spaces are separable. Rather, I would present it as probably the most basic example of a non-separable space. The important point seems to be that there was no particular reason to think that all Banach spaces would be separable, and finding an example of a non-separable space is not very difficult.

I think the first point is more important here: we are more inclined to call an object a counterexample if the existence of that object disproves a statement that we had quite good reason to believe. It often happens that after repeated unsuccessful attempts to prove a statement, mathematicians begin to feel that it has no particular reason to be true, even if it seems to be hard to come up with a counterexample to it. In such a situation, if a counterexample is eventually found, it may have lost something of its “counter” feel. My impression is that the construction of a non-sofic group comes into this category. There have been several proposals in the literature for how one might construct such a group, and I don’t think there were many (or even any?) experts who strongly believed that all groups were sofic. So it feels more natural to say, “OpenAI came up with the first example of a non-sofic group” than to say, “OpenAI found a counterexample to the soficity conjecture” (despite the fact that that section of their paper is entitled “A counterexample to the soficity conjecture”).

Likewise, it seems to me that the new lower bound for multicolour Ramsey numbers is more of an example than a counterexample. I think quite a lot of people believed that the bound should be exponential, so for them it was a counterexample, but others, myself included, were more neutral about it. As a matter of fact, I have worked on the problem in the past (a long time ago) in an equivalent formulation, which asks how many triangle-free graphs on n vertices you need if you want their union to be the complete graph K_n. If you take bipartite graphs, then it’s easy to see that you need \log_2n of them, but that bound can be improved if instead you observe that a complete 5-partite graph can be written as a union of two triangle-free subgraphs, and therefore it is possible to write the complete graph as a union of 2\log_5n triangle-free graphs. It is then tempting to try to do better, with triangle-free graphs that are less dense but that make up for it with unbounded chromatic number — a necessary condition if one wishes to use a sublogarithmic number of graphs, which is equivalent to showing a superexponential lower bound for R(3,3,\dots,3). All this is to say that when I worked on the problem, my efforts were concentrated on what turned out to be the right direction, so for me OpenAI found an example of what I (weakly) expected, rather than a counterexample.

Where does this leave us?

I would like to find a coherent explanation of the conjunction of the following facts.

  1. The most notable mathematical results proved by LLMs have tended to be ones that we would classify as examples or counterexamples, where counterexamples are, broadly speaking, existence statements that disprove statements that we expected to be true.
  2. Many statements can be formulated as existence statements when we would usually think of them as universal statements, and vice versa, so what we consider to be an example depends on the mathematical context of a statement as well as its logical form.
  3. LLMs are pretty good at proving universal statements as well: it’s just that the strongest statements they have proved that we would think of as theorems have mainly not been at the level of the strongest statements that we would think of as counterexamples.

Given these facts, it seems likely that what LLMs are good at is something else, which happens to have as a consequence that they are good at the kind of existence problem that we would normally classify as asking to find a non-trivial example.

Let us consider two things that we can be confident that LLMs are good at. One of them is knowing a lot of mathematics: if a problem can be solved by means of a relatively standard argument, it is highly likely that an LLM will be able to find and use that argument. The other is the ability that an LLM has simply by virtue of being a computer: it can work at huge speed (compared with humans at least) and can therefore afford to make a large number of unsuccessful attempts at a problem before it finds a solution.

Without even looking at what LLMs have actually managed to solve, one might guess that these two features would lead to their having a somewhat different style from human mathematicians. Very roughly, LLMs would have the edge when there is more of a probabilistic element to the proof-finding process: they would be good at problems for which the best method is to try a lot of ideas, not necessarily particularly novel, until at some point you get lucky. Humans on the other hand would be better (for the moment) at finding more “surprising” and “conceptual” arguments, where the appropriate method is to dig deeper and deeper into a problem until the solution reveals itself. (It is hard to say exactly what this means, but I hope that any experienced researcher reading this will know what I am talking about.)

This raises two questions: does the guess above correspond at all to the reality that we are observing, and is there any reason to suppose that what I have tentatively described as the “LLM style” of doing mathematics would lead naturally to LLMs discovering several counterexamples (or just examples) to long-standing conjectures, even if that was by no means all they could do?

I don’t pretend to have a scientific answer to either question, but the reactions of experts to several of the remarkable solutions that ChatGPT has found do lend some support to the idea that LLMs work in more of a try-lots-of-things-till-you-get-lucky way. People often seem to react by saying something like, “Initially I was amazed that the problem had been solved, but on closer inspection I realized that the approach was actually not all that novel, and one that with the right small hint a suitably expert human could have found quite easily.”

For the second question — whether the LLM style is well suited to finding (counter)examples — I think matters are less clear, because there are many ways of searching for a counterexample, and some of them fit better than others the style I have described. Here are a few general methods. (I don’t claim that the list is exhaustive.)

  1. Look for an off-the-shelf example. Here one has a stock of fairly standard examples and one simply tries them out one after another to see whether any of them fails to satisfy the given statement. For example, Ryan O’Donnell ends his wonderful book on the analysis of Boolean functions with some tips, one of which is, “If you have a conjecture about Boolean functions, test it on dictators, majority, parity, tribes (and maybe recursive majority of 3). If it’s true for these functions, it’s probably true.”
  2. Build an example from basic examples and standard construction methods. For an algebraic problem, for instance, one might start with some standard examples, but then take products or quotients or limits.
  3. Make heavy use of metavariables. The word “metavariable” comes from computer science, and in particular from automatic theorem proving, and refers to the practice that in mathematics would correspond to writing, “where x is to be chosen later,” (in which case x is the metavariable). In a paper we usually do this only in fairly simple situations such as when we need to choose a number \epsilon>0 that is small enough for later arguments to work. But when we search for an example of an object x that satisfies some property Q (which may well be a conjunction of simpler properties Q_1,\dots,Q_k), it is often not a good strategy to specify x completely and only then to check whether it satisfies Q. Instead, it can be more fruitful to do almost the opposite: we start by saying virtually nothing about x and simply launch into proving that it satisfies Q. In the course of doing so, we find that we need x to satisfy a property P_1. If we are lucky we can describe in a nice way a very general class of objects x that satisfy P_1. For instance, we may be able to find a parametrized class: we identify some function f and show that f(y) satisfies P_1 for every y of a certain type. The problem is then reduced to finding y such that $Q(f(y))$ holds, which is a more specific version of the original problem. There may be many iterations of this process, or a mixture of this process and other processes, before an example is eventually found.
  4. Try to prove the opposite. If one wishes to find x such that Q(x), it can be surprisingly helpful to start by attempting to prove the statement \forall x\ \neg Q(x). The reason this can be helpful is that using our standard methods of attempting to prove something, we may end up identifying a key lemma that would suffice: that is, we may find an intermediate property R that implies \neg Q in a non-trivial way and thus reduce the problem \forall x\ \neg Q(x) to \forall x\ R(x). Turning things round again, it may well then be that finding a counterexample to R is easier than finding a counterexample to \neg Q (that is, an example that satisfies Q). Of course, there is no guarantee that a counterexample to R will be an example of Q, but sometimes we are lucky and it is. More often, we can use the idea of the previous method, noting that it is at least a necessary condition of an example of Q that it should not be an example of R, so one can try to describe a general class of objects that fail R and in that way reduce the problem.
  5. Successive approximation. Sometimes, when we are searching for an example of x such that Q(x), we write down a moderately plausible guess x_0 not because we think it has a chance of working (if we did, then we would be using the first strategy), but because we hope that if x_0 does not satisfy Q, then we will be able to diagnose what went wrong and specify a new guess x_1 that does not have that defect. Again, this strategy can either be iterated or combined with one or more of the other strategies.
  6. Just-do-it proofs. Sometimes we need x to satisfy infinitely many properties Q_1,Q_2,\dots, each of which is, individually, quite easy to satisfy. In such situations, we often “build” x inductively bit by bit, ensuring at the ith stage of the process that however the building process continues, x will satisfy Q_i.
  7. Pick a random example. Often it is very hard to give an explicit example of an x that satisfies Q, but there is a natural probability distribution for which one can show that if one chooses x randomly from that distribution, then with high probability (or at least non-zero probability) it will satisfy Q.
  8. Pick a generic example. In more infinite contexts, it may again be quite hard to give an explicit example of an x that satisfies Q, but one may be able to show that the set of x that fail Q is or measure zero, or is a meagre set, or is small in some other way.

There is no particular reason to suppose that LLMs would be equally good at each of the methods above. So perhaps what we are observing is not quite that LLMs have a particular ability to find examples, but more that they are particularly good at finding examples (and proofs) in a certain way. Looking at the above techniques, one might imagine that they would be very well suited to checking off-the-shelf examples, finding just-do-it proofs (since that is a rather standard method with lots of instances in their training data), using the probabilistic method (unless, as often happens, significant new ideas are needed to show that the probabilities work out), and picking generic examples. The other three methods described above — use of metavariables, trying to prove the opposite, and using successive approximation — require more of an ability to judge whether the approach one is taking is likely to be fruitful. Here it seems at least possible that humans will sometimes have an advantage, but the conditions that a problem would need to satisfy are quite stringent. One would need an example to be one that lies at a leaf of a very large search tree — too large to be searched for by a combination of moderate mathematical ability and brute force — but that can be found by a mathematician with a sufficiently good nose for when they are making progress that they can prune the search tree very substantially.

Why wouldn’t LLMs also have that “nose”? I don’t rule out that “nose” is an emergent property of the way LLMs are trained, and that within a year or two they will have it to the same extent that we have it. But for now, in my interactions with ChatGPT, I do have a distinct impression that they haven’t got there quite yet. When I discuss an open problem with 5.6 Pro, I am often presented with approaches that sound promising until I think about them carefully, and then seem quite a lot less promising. And they will also often end a response by saying, “I have not managed to answer the question you asked, but have managed to reduce it to the following much narrower and more precise question,” which sounds very promising until it has happened five times without any obvious progress having been made. It isn’t completely obvious how they will get better at this, since their training data will not be full of examples of fruitful and less fruitful directions to pursue when trying to solve problems: all they will typically see is tidied up proofs that hide the thought processes of their discoverers. Of course, human mathematicians also don’t get to learn much about how to do research from the experience of other mathematicians, and yet we somehow manage to pick it up. But the situation is a little different for us, in that a lot of what we learn is by doing rather than emulating.

Another reason it is not obvious that “nose” is a property that emerges naturally when LLMs are scaled up is that if LLMs make heavy use of their broad knowledge and can afford to do a lot more brute-force search than humans can, then they will lack the incentive that humans have to prune the search tree ruthlessly. It could conceivably be that their successes so far are achieved using methods that for a human would be considered extremely inefficient, but that because of their superior speed and knowledge, the combinatorial explosion these methods will lead to has not yet become apparent.

It would be very interesting to try to test this experimentally, but it is also difficult, because if an LLM has what looks like the kind of idea that could only be the result of “deep thought” about a problem, we can never be sure that it has actually carried out that deep thought, as opposed to finding a model argument already in the literature, or in other words exploiting the deep thought of a human mathematician. It would probably be easier (but still not easy) to test it by using models that are less powerful than the latest ones and that have been to some extent shielded from the mathematical literature: one could give them a carefully designed suite of problems and see whether the ones that the LLMs solve have particular characteristics.

It may seem as though I am desperately clinging to the hope that humans will continue to be able to make meaningful contributions to mathematical discovery for a while yet, but while I do indeed hope that, I am not making any assertions of the form “LLMs will never be able to do X”. I think it is likely that they will, and given the pace of progress over the last three years it will probably happen quite soon. But I do think that there may be a hurdle for LLMs to clear and it seems at least possible that it won’t be cleared as straightforwardly as some of the previous hurdles.

In that connection, it would also be interesting to see whether a different reward structure leads to LLMs being able to solve different kinds of problems. For example, if during training an LLM (or machine-learning system of some other kind) is not just rewarded if it ends up with a solution, but also penalized if it explores too many dead ends or if it “cheats” by getting the answer from the literature, perhaps it would be incentivized to go about the research process in a more human way and thereby achieve better results for classes of problems where it is yet to make a big impact.

If the hurdle is cleared, either by pure scaling up or by some more thoughtful method, it will be quite difficult to know when that has happened, since, as just mentioned, an idea that seems very original and surprising may just be lurking somewhere in an LLM’s training data. But I would be confident that it had been cleared if an LLM were to come up with a proof that was as surprising to me as the solution of the cap-set problem was in 2016: the previous best known bounds were completely eclipsed, the method was utterly different from anything I had thought about trying, and afterwards there was a flurry of activity as people came to understand what this wonderful new technique was capable of.

Conclusion

I wasn’t quite sure where I would end up when I started this post, and now that I’ve got to the end, I feel that my main conclusions are not particularly new or surprising, but I hope that the route to them is of some interest. The main points I have made are the following.

  1. “Finding an example” is in practice not the same thing as proving a statement that begins with an existential quantifier.
  2. If it is true that current models are particularly good at finding examples, that is probably not because they have a particular affinity for existential statements, but more because the proof-discovery methods that are appropriate for finding certain kinds of examples play to the obvious strengths of LLMs: wide knowledge and the ability to explore many paths of the search tree that humans would judge to have a low probability of success.
  3. It seems likely that LLMs will carry on improving very quickly. However, if, contrary to expectations (mine at least), there turns out to be some residual class of problems (or other mathematical activities) for which humans continue to have the edge for a while, it is likely that those will be problems for which the mysterious human ability to prune the proof-discovery search tree is particularly advantageous: that is to say, problems where the search tree is deep and has a large amount of branching, so that without rigorous pruning a search is not feasible even for a computer.
  4. A good sign that LLMs have reached human level for a much wider class of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but that with hindsight come to seem beautiful and natural. They should also be methods that are difficult to stumble on by accident. It is hard to say precisely what would count as such a proof, but I think we’ll recognise it when we see it.

July 29, 2026

Secret Blogging Seminar — An experiment with AI-assisted writing

As in David’s most recent post, there’s been a lot in the news about finding proofs and counterexamples with AI. Last weekend, I decided to try an experiment with writing using AI. I learned a lot, and wanted to quickly discuss the experiment and my thoughts on it here. Lots of people are certainly already doing this, but I haven’t seen many people talking about it.

The starting point is that Victor Ostrik and I started a project back in 2017, generalizing a result of Kuperberg about quantum G2, from generic q to q a root of unity. Namely, we showed that for q a root of unity outside of a specific finite list, the Karoubi completion of the G2 spider category is equivalent to the category of tilting modules of the Lusztig form of the quantum group G2. At some point during those 9 years, we did a little bit of writing, and at some point I gave a talk on it, but otherwise we did very little writing. This was not for mathematical reasons, but rather for executive function reasons on my end, the global pandemic, and both of us becoming directors of graduate study. This suggested an interesting challenge: could I use LLMs (specifically ChatGPT 5.6 Sol work mode mostly at “very high” intensity, via IU’s “Edu” subscription) to write this paper that was essentially mathematically complete, but almost entirely unwritten, and how quickly could this be done. To some extent this was a free experiment, because realistically I don’t think we’d have ever finished the paper at this point, and so it’s not replacing a bespoke paper that could have existed.

After spending a decent chunk of the time from Saturday until now on it, I now have a draft that I’m pretty happy with. I want to emphasize although mathematically this is Victor and my joint work, and although Victor has allowed me to make this post, he has not signed off on the accuracy and all errors at this point should be blamed entirely on me. Also my work is supported under NSF DMS grant 2000093 and Simons Foundation grant MPS-TSM-00007608.

Ok, here’s what I did:

  1. First, I asked if Sol could one-shot the main theorem. The answer was yes, though for a somewhat simple reason: Bodish-Wu write “It is possible to adapt the approach from [1], which itself is based on [7], to prove that the Karoubi envelope of [the G2 web category] is equivalent to the category of tilting modules as long as $[2], [3] \neq 0$.” That is to say, Elijah already proved the same result for C2, and a similar argument will work for G2. So the robot supplied the similar argument. I asked it to write that argument up, and then to check it over for good references and to read it like a referee would and make edits. This took around 30 minutes. Here’s the resulting file.
  2. Second, I uploaded my talk slides (and the tiny file already written, which was mostly useless), and asked Sol to give a proof of the main results following the slides. Again I asked it to edit it. This took around 30 minutes. Here’s the resulting file.
  3. Then I looked at the files. As mathematical exposition, I consider both to be garbage.
  4. Then I spent several days giving feedback attempting to improve the second file based on my talk. At no point did I edit the source directly. Most of this was in what I would call the style of a (low executive function, see above) PhD advisor. That is, I would kinda skim the file, get annoyed about something, and tell it to fix it. While it was fixing the paper, I would skim some more to try to find something else that annoyed me. This was a long process! It took three days, nearly 100 prompts, 10-15 hours of reasoning, plus another 10-15 hours of non-reasoning computer time. This used nearly an entire week of my generous budget, and Sol estimates that this would cost around $100 (within a factor of 2) at metered rates. Eventually I got to a version of the paper that I’m pretty happy with. Here’s the resulting file.

I thought I’d distill some thoughts and some questions from the process, I’m of course very curious for your thoughts on the matter.

Comments:

  1. This was much faster than I could have written the paper myself, though slower than I thought it would be. I think the final product is comparable in quality to a typical math paper of mine. On the other hand, I think that compared to my fastest writing collaborators it was not orders of magnitude faster, and the quality is not close to the output of the best mathematical expositors. AI at this point is much worse at writing paper than finding counterexamples to conjectures.
  2. In this case, I was not very worried about errors, because I already had thought through the whole argument and was highly confident that it would work (modulo getting the exactly correct list of exceptions). Nonetheless, I felt like Sol did not make errors more frequently (or of a worse character) than I would expect of myself or a collaborator. Most errors were stuff like “Oh, forgot to check whether this theorem actually works at all roots of unity.” This is typical of my experience with 5.6, which is dramatically better at doing math accurately than previous ChatGPT models.
  3. In this case the vast majority of the ideas were already present from Victor and my work. In particular, the goal was not just to write a proof, but to write our specific proof. Nonetheless, I do think the model contributed mathematically in one key way: in my original sketch I always worked over each q individually, and the model preferred to work integrally, and this resulted in some very nice simplifications in Section 4.1. If and when we turn this into a real preprint, I will include a brief discussion of the intellectual contribution from the model.
  4. I was surprised when I printed out and read a near-final draft, that this feels to me like a paper I wrote. That is the voice is not different enough from what I would write with a human collaborator to feel like it’s not in large part mine.
  5. The experience is disconcertingly similar to advising a PhD student on a paper. That said, a PhD student would need less handholding on their second paper, but an LLM won’t really learn.
  6. I was surprised about how important “prompt engineering” remains, and I think that if I were to write another paper this way I would be able to write it faster and better. The key points are that the model is lazy and easily distracted (both properties I find highly relatable!). It’s lazy in the sense that if you ask it to do a lot of work all at once it will take shortcuts and not do a good job. At one point I had to be like “no, go look at exactly how I made TikZ diagrams, now make all your diagrams actually good like that.” It’s easily distractible in that if you’re not clear about the scope of your question and the document is long, it will start spending crazy amounts of time doing who knows what. Like it wrote the whole first draft in 20 minutes, but then when the paper was 50 pages long, I asked it to switch the order of two paragraphs and it took an hour. Make clear requests and not too many requests at once. Form a plan first and then implement the plan. Be specific about whether it should be editing the document, and if so in which sections. For simple tasks, medium intensity is better than very high.
  7. Starting again from sketch, I’d try to follow Terry Tao’s advice for writing and start with an outline and gradually flesh it out, rather than trying to start with a one-shot paper and then editing.

Questions:

  1. To what extent is this final paper adding any value to the original talk? Especially considering that readers themselves could use an AI model to flesh out points in the talk that they didn’t understand? Maybe we should just be focusing on talk-length digests and formal checking, rather than traditional papers?
  2. What should we do with this paper? I don’t want to make someone hand-referee it, because it doesn’t seem fair when it wasn’t hand-written. Probably we will put it on the arxiv once we’ve human-checked it fully and Victor has signed off on it, so that other people can use the results if they need to.
  3. Given the speed-up, when does it still make sense for me to write papers by hand? (Relevant here that I’m a very slow writer and don’t really enjoy it, the way I enjoy say preparing and giving a talk.)
  4. What does this mean for PhD advising? Many PhD students need a similar amount of guidance to what I gave the model in this project. But you can now remove the student from the loop (either intentionally, with the advisor just writing using LLM assistance rather than having students, or unintentionally, with the student just feeding all the suggestions to an LLM and reporting back to the advisor).
  5. Have any of you done better with AI-assisted paper writing? My points 6 and 7 above sounds like something where someone is going to say “blah, blah, scaffolding, blah, blah, multi-agent…”

What a strange world to live in…

July 26, 2026

Tim Gowers — Thoughts about the Leiden Declaration

Last September I went to a workshop at the Lorentz Centre in Leiden to discuss mathematics and AI with historians, philosophers, computer scientists, AI researchers, and mathematicians of several different flavours (though there was a surprising preponderance of algebraic geometers). The whole event was extremely stimulating, with some talks but also a lot of time set aside for discussion. One of the concrete outcomes of the workshop was the Leiden Declaration, which has now been signed by over 3000 people. Given that I was part of the workshop, it might seem a bit strange that I am not one of the signatories of the resulting declaration. The reason is not so much that I disagree with it in any concrete way, but more that in several places it makes confident assertions and recommendations that I feel somewhat uncertain about. So instead I prefer to try to articulate my views about the issues raised by the declaration and put them in this blog post. Before I do that, I would like to make clear that I am very glad that the Leiden Declaration exists and I think that it has done a lot of good in focusing people’s minds on the issues that AI is forcing the mathematical community to grapple with, which are more acute now than they were last September.

Let me begin by quoting a passage from the declaration that sets out “what we take to be characteristic values of mathematical research that we have a joint interest in preserving”.

  1. There are many reasons to pursue mathematical research, ranging from intellectual curiosity to a desire to solve practical and societal problems. Underlying much of mathematics is the activity of proof. Mathematical proofs are regarded as conferring the highest degree of certainty to their conclusions, as well as imparting understanding of why their conclusions are true. These characteristics of proof support the scientific integrity of mathematics.
  2. Results are attributable to specific authors who take credit for their discovery and assume responsibility for their correctness. These principles ground the merit-based standards to which we aspire in mathematical research.
  3. Mathematical arguments are regarded as transparent and subject to independent verification. They may be extremely long or difficult, but in principle no proprietary knowledge or equipment should be required to understand them.
  4. Mathematicians share a concern for proper evaluation of mathematical work relative to shared standards of depth, difficulty, and significance.
  5. Mathematics produces not only a body of results, but also understanding, clarity, and judgment among the communities of mathematicians who have shaped them, often in the context of their own autonomously guided research. This expert knowledge is essential, both to effectively use mathematics, and to continue to articulate new and significant research questions. A key source of strength of the discipline has long been the autonomous shaping of the direction of research and the methods used to pursue it.

The first thing I would say about these values is that they are undoubtedly values that are widely held by mathematicians, including, with some qualifications, me. The main qualification I have concerns point 4: I find the notion of “proper evaluation” somewhat problematic, given that different mathematicians can have very different judgments without either of them being clearly wrong, especially when it comes to the significance of a piece of mathematics. Also, these judgments are used for purposes such as the acceptance of papers in journals, hiring and promotion decisions, the awarding of prizes, and so on, that are part of a system that copiously rewards a few people — I myself have hugely benefited from it — but doesn’t necessarily adequately reward a lot of people who are doing less visible work that is essential to keeping the whole enterprise going.

But the more important point is whether these values are ones that we should fight for in the future, as the Leiden Declaration suggests. I find that clearer for some of them than others. For example, it seems to me that the importance of rigorous proof will be even greater in an AI age than it was before — if the output of AI is not underpinned by rigorous proof, then the kinds of difficulties one already hears about with certain areas of human mathematics (see for example many talks by Kevin Buzzard arguing for the value of formalization) would be hugely magnified. But what about the attribution of results to specific authors, who take both credit and responsibility for them? Suppose that at some point in the future AI becomes more autonomous, reading the literature and solving many problems that it finds. Suppose also that its solutions are autoformalized, so there is no serious doubt about their correctness. In such a situation, there would be nothing for a human to take credit for or responsibility for. Does that mean that we should declare such results undesirable and threatening to mathematical values?

Of course, something could well be missing in such a situation: perhaps the proofs would be badly written and hard to follow, which would mean that they lacked something we all very much value. So let me extend the thought experiment slightly. What if by that stage one could take one of these outputs and ask an LLM to explain the ideas, and what if LLMs did a very good job at that? That is not particularly hypothetical, since they are often pretty good at this job already, but I am imagining a world in which they are much better than they are now, as they will presumably become.

So now we would have a world in which a lot of problems had been solved, we were sure that the solutions were correct, and we had an LLM ready to explain those solutions in as much or as little detail as we wanted. Is that a future we should resist, and if so, why?

One obvious reason is that it would take a huge part of the fun out of the subject. It is extremely satisfying to struggle with a mathematical problem for months or even years and eventually solve it. But I worry about that argument, because it seems to be saying that we should resist doing mathematics the easy way because a tiny fraction of the world’s population gets huge pleasure from taking orders of magnitude longer to do it. That is not to say that I wouldn’t be sad that a way of life that has sustained me for the last forty years was not available any more — of course I would. I just find it hard to use it as a reason to argue that we should try to preserve the “ownership structure” of mathematical results. If we arrive at a world where mathematical theorems are no longer associated with mathematicians, maybe that won’t be any more problematic than the fact that stars aren’t named after astronomers and most aren’t named at all. I’m not necessarily in a hurry for that world to exist, but maybe once the transition had happened, people would be OK with it.

The third value I share in an uncomplicated way, and I have already discussed the fourth. The fifth value is one that I hold very strongly, though I’m not so keen on the idea of experts consciously “shaping the direction of research”, something that I see as happening more organically. Obviously there are some notable examples of mathematicians who have created wonderful programmes of research, but even there I would like to credit other mathematicians with understanding what is wonderful about those programmes and contributing to them enthusiastically as a result, rather than being told what direction to pursue and meekly doing so (which is probably not what the declaration is actually trying to suggest, but it has a slight flavour of that for me).

But that’s a minor quibble when set against my main worry about the effect of AI on mathematics, which is the possible destruction of mathematical culture. There is at the moment an extraordinary body of knowledge and expertise that exists not just in the mathematical literature but in the heads of mathematicians all round the world. Imagine if AI didn’t exist and a pandemic broke out that for some reason wiped out all mathematicians and nobody else. All the literature would still be there, but nobody would have the faintest idea what to do with it. To revive a mathematical tradition under those circumstances would be extremely difficult and take decades. Now imagine a slight variant of that, where AI does exist and because of it people are no longer motivated to put in the years of effort it takes to reach the level of expertise that a typical research mathematician has now. After a decade or two, we might arrive at a situation where the mathematical literature has, in some form, been vastly expanded, but there is no corresponding community of human experts who have a shared understanding of parts of it. Almost all of mathematics would be like the areas that we have more or less forgotten about today, areas that exist in papers written many decades ago that nobody reads any more. (I won’t name any such area because I don’t want accidentally to suggest an area that many people still love and work on.)

This, it seems to me, is a possibility that we should try very hard to resist, but I agree with many other commentators who say that in order to resist it, we will need to give less priority to some of our current values — and I would include ownership of mathematical results in that list — and more to others. For example, if Person A gets an LLM to one-shot a solution of an important open problem (which is formalized, possibly automatically, so there is no doubt about its correctness) but Person B makes the effort to digest the solution and explain it in a way that other mathematicians can understand and learn from, then I think we will want Person B to get the lion’s share of the credit. The credit would be of a slightly different from what it is now, which could be described as admiration for somebody’s talent, insight, speed (I mean here the purely factual statement that speed is often admired — I would prefer that to be less the case) and hard work. It would be more like the gratitude that one feels already for somebody who writes a beautiful textbook that makes a whole area of mathematics coherent and accessible.

Maybe that is what the “research mathematicians” of the future should do: make a selection from a vast sea of AI-generated mathematics and write a book about it in such a way that other mathematicians can read the book and feel the kind of enrichment that we feel when we get to grips with an area of mathematics.

At this point I have to admit that there’s a pessimistic side of me that asks the following general question whenever anyone says anything about what the role for humans might be in the future: why do you think that AI wouldn’t be able to do it? For example, with the suggestion I’ve just made, what reason is there to suppose that ChatGPT 8.2 wouldn’t be able to have a short interaction with you about your mathematical tastes and background and then write the ideal textbook just for you? Humans are likely to be better at this kind of curating for a little while yet, but is it a fundamentally human ability that AI could never hope to emulate?

In a world where AI wrote bespoke textbooks (or more likely, just taught people in some more direct way), something would be lost that feels important: mathematics as a collective endeavour. If we all just learnt cool bits of maths for our own private satisfaction, we would miss the considerable pleasure that comes from discussing mathematics with others, though even that could in principle be restored by a benign LLM that deliberately taught many people the same cool bits of the subject, though an LLM that could do that sort of social engineering would raise all sorts of safety issues.

Let me now turn to the section of the declaration about potential threats. I’ll put my comments on each one in square brackets.

  1. Current automated techniques can produce plausible but unreliable (or even incorrect) arguments which are difficult to distinguish from correct mathematical proofs. This applies not only to informal arguments, but also to formalizations, where the difficulty lies in the translation between computer-encoded and human presentations of concepts. These fast-moving developments put our present system of review under increasing pressure, jeopardizing our ability to implement traditional standards for the correctness, transparency, and independent verifiability of proof. [This feels like less of a problem now than it did last September, partly because the best LLMs hallucinate a lot less than before, and partly because autoformalization is improving all the time — I have just used harmonic.fun’s Aristotle system to formalize a complicated paper in Lean and I didn’t need to know any Lean to do it.]
  2. Technologies that draw extensively on the published mathematical commons undermine the traditional system of attribution. Models trained on published works frequently return outputs that do not properly cite the human works they synthesize. Many current models are also built on data obtained by systematically exploiting licenses and access arrangements that were not made with artificial intelligence in mind, or indeed by simply violating copyright protections. [This is a problem at the moment, when ownership of results is important, and I am very much in favour of people making an effort to give appropriate credit for mathematical ideas that AI may have used. However, in the longer term, as I have already discussed, I think this ownership structure will break down and the issue will become less important. It also seems possible that LLMs will become better at revealing their sources.]
  3. Technologies which affect the way in which mathematics is practiced may disturb the current system of incentives. The use of artificial intelligence — and thus also the sort of problems which it can address — may become incentivized for its own sake, disrupting our mechanisms for hiring, funding, and recognition. This disadvantages researchers who do not have access to the technologies or decision-making related to them, or who are unwilling to use technologies controlled by organizations whose values they do not share. [These seem to me to be genuine problems. I think there is simply no point in hoping that our current system of incentives will not be disturbed — it obviously will. I am not necessarily too worried if our mechanisms for hiring, funding and recognition are disrupted, as I don’t find those mechanisms unproblematic as they are, but disadvantaging researchers who do not have access to good LLMs is something I certainly think we should worry about.]
  4. Proper evaluation is endangered if results are communicated through informal channels such as press releases or blog posts, often without any research paper or other disclosure of information necessary for scientific evaluation. This practice seeks publicity for new results on market timelines before the accepted processes of community evaluation in mathematics can take place. In many cases this leads to simplifications in reporting, such as overemphasizing the significance of automated tools and undervaluing the prior human contributions which have made those tools possible. Such oversimplification risks influencing public opinion in a way that not only damages perceptions of mathematics, but also misleadingly uses specific mathematical tasks as metrics for the general reasoning capacities of commercial products. [I think this can be a problem, but I think it is not as serious a problem as some of the others, since when results get overhyped, there seems to be no shortage of people publicly (and rightly) pointing that out.]
  5. These developments put the autonomy of mathematics under threat. The increasing involvement of technology companies in mathematical research raises the risk that research questions may come to be prioritized because of their amenability to automated mathematics, rather than expert judgment of their deeper significance. Indeed, broader understanding of the field may be permanently lost in the process of automation. With university budgets under pressure, this reshaping also changes professional incentives in a manner which encourages the collaboration of researchers with technology companies on asymmetric terms. If left unchecked, these trends go beyond threatening researchers’ autonomy, affecting the scope and depth of mathematical research itself. [I think this could be a problem, but it also seems to me that mathematicians have a lot of power here. For instance, if a technology company were to produce a lot of research that mathematicians did not find all that interesting or important, I don’t think they would be able to use their financial and other resources to persuade us to change our minds. Rather, what seems to happen is that mathematicians say, “Yes that does X but it doesn’t do Y,” and the tech companies then feel challenged to do Y.]

There follow eleven recommendations for individual mathematicians. I agree with almost all of them. The one that I’m not so sure about, for reasons I’ve basically already gone into, is this.

Affirm the humanity of authorship. Credit and responsibility continue to belong to humans within the mathematical community and should not be given to automated systems. Artificial intelligence may obscure, but does not replace, the collective human labor behind a result.

I’m not sure what that really means. For example, should we affirm the humanity of authorship in the case of the solution to the unit-distance problem? Some humans did a wonderful job of explaining the proof that OpenAI’s model came up with, and the model made use of some highly non-trivial mathematics produced by humans, but the solution itself has not been credited to any human, and nor should it be in my view.

Under recommendations for mathematical organizations and not-for-profit research funders I again agree with several of them but have my doubts about some. An interesting case is the following.

Protect the rights of authors. Automated mathematics presents new challenges to the rights of authors, and societies should be proactive in the development of sample licensing agreements to protect these rights. In particular, material should not be used as training data without consent, and publishing agreements should allow authors to opt-out [sic] of the use of their work in this way.

This recommendation seems to belong to a world in which journal articles are the main means of dissemination of mathematics. But that has long since ceased to be the case: almost all dissemination now takes place via arXiv preprints, with journals limited to providing a little extra mark of prestige. Once an article is on arXiv, it is on the internet and one can hardly ask for it not to be used as training data. So this recommendation, if it applies at all, will apply to a tiny fraction of articles that are published without first appearing on arXiv. More generally, what right of an author is being compromised when an article is used as training data? We don’t object if human mathematicians use our articles to help train themselves to become better mathematicians — indeed, we will typically be delighted that somebody else thought our articles worthy of their attention. So the objection to a machine doing the same would have to be that for some reason one did not want machines to get better at mathematics in a similar way. I can imagine grounds for such a wish: perhaps somebody is worried about the threat that LLMs pose to traditional mathematical practice, or perhaps they worry that mathematical ability of LLMs will transfer to much more dangerous reasoning ability. But there’s a more complicated discussion to be had here than one might think from reading the recommendation.

The next recommendation is this.

Insist on appropriate publication outlets. Demand that mathematical results continue to be published in peer-reviewed venues such as journals, proceedings, and books. Informal mechanisms such as press releases or blog posts can provide a valuable supporting role, but they cannot replace peer-review or community scrutiny.

For reasons that I’ve gone into many times, I am not too fond of the current publication system, so I can’t get behind this recommendation. Indeed, if the current system becomes unsustainable because of a flood of AI-generated and AI-aided content, I would regard that as a beneficial consequence of AI. However, that doesn’t mean that I would advocate a total free-for-all. I’ve already said that one of my worries is that if mathematical content is not sufficiently organized, then the traditions that we all value could die. I just think that what we will want to do to preserve those traditions is likely to be a lot more innovative than clinging on to the peer-reviewed journal system.

I have highlighted in this post the parts of the declaration that I have doubts about, either because I disagree with them or, more typically, because I sort of half agree with them but want to add many qualifications. That may make the post come across as rather negative, but that is not my intention. The parts I disagree with are in the minority, and I think it is important that a declaration such as this should be made. I should also make clear that my views are evolving all the time, largely because the speed of progress of LLMs has taken me by surprise, but also as a result of conversations I have had or opinions that other mathematicians have expressed online.

I’ll end with two further clarifications. The first is that it may seem as though I am taking it for granted that LLMs will soon be better than humans at all aspects of mathematical problem solving, and maybe also problem posing, theory building, formulation of definitions, etc. I do think all that will happen at some point, but whereas some people say that it will obviously happen within the next two to three years, I would say that it might happen as soon as that, but I don’t rule out that we’ll get lucky and find that we can do interesting AI-assisted maths for quite a bit longer than that before AI doesn’t need us any more.

The second is that I think I have acquired a reputation as somebody who celebrates what is going on. But if, for example, I post on Twitter saying that such-and-such an AI solution is a remarkable development, the word “remarkable” is meant to indicate no more nor less than that I found it very surprising. My feelings about the possibility of AI solving all sorts of problems that interest me are much more mixed. I’ve had the experience twice now of seeing GPT 5.6 Pro one-shot a solution to a problem that I very much liked and had thought about hard (in both cases with much younger collaborators, who, with my approval, were the ones who prompted the LLM). It felt very strange and not particularly pleasant to have the rug pulled out from under my feet like that. On the other hand, I was quite pleased to see the problems solved. It’s actually a similar feeling to the one I have had many times when a problem I am fond of and have thought about gets solved by another human mathematician.

Another factor for me is that I have invested a lot of thought into automatic theorem proving of a more traditional kind. One of my main motivations for that was the hope that the work I put into it would extend the state of the art, measured by which problems a computer can solve. That ship has sailed now, and that saddens me. I still think that there is value in the work that I and my group are doing, but it has become a tougher sell.

So I personally have already found AI quite disruptive, and this is just the beginning. I would have preferred the developments to happen at a slower pace. But I don’t see any practical way to slow them down, so the best we can do is probably to face up to the changes that are being thrust upon us and do what we can to maximize the benefits and minimize the damage. The Leiden Declaration may not be perfect, but it makes an important and positive contribution to that effort.

July 25, 2026

Clifford Johnson — On top of the Mountain again

Just in case you’re up for a short talk at the top of Mount Wilson followed by an evening of observing through the historic telescopes on Saturday 25th July… this might be for you! Go to Mount Wilson Observatory’s website for more. –cvj

The post On top of the Mountain again appeared first on Asymptotia.

July 24, 2026

Peter Rohde Introducing Sigfried’s Blog

My new secondary blog featuring conversations with AI, inventing new things, exploring hypotheticals, letting creativity flow freely.

Some highlights:

  • Satellite constellations with topologically distributed apertures.
  • A clockless architecture for classical topological computing.
  • Post-quantum cryptography using the \mathbb{Z}_2^n \rtimes S_n algebra.
  • Efficient homomorphic computing using reversible classical circuits.
  • A silent speech interface using microwave Doppler imaging.
  • Cognitive search acceleration.
  • Consensual thought guidance.
  • Subliminal audio modulation & human guidance systems.
  • Microwave imaging using WiFi and 5G for medical applications.
  • Thought tomography.
  • The quantum bluff hypothesis.

https://sigfriedschattenjaeger.wordpress.com

July 20, 2026

Secret Blogging Seminar — The new counterexample to the Jacobian conjecture

As many of you have probably heard already, yesterday morning, Levent Alpöge tweeted that Fable had found a counterexample to the Jacobian Conjecture. Specifically, let

a=(1+xy)3z+y2(1+xy)(4+3xy),b=y+3x(1+xy)2z+3xy2(4+3xy),c=2x−3x2y−x3z,\begin{align*} a&=&(1+xy)^3z+y^2(1+xy)(4+3xy),\\ b&=&y+3x(1+xy)^2z+3xy^2(4+3xy),\\ c&=&2x-3x^2y-x^3z, \end{align*}

Then the Jacobian of (a,b,c) is easily checked to be -2. However, the map (a,b,c) is generically three to one, not bijective.

I’m sure many of you are playing with these polynomials to see what you can figure out about them. This is a place for us to share our observations. I’ll post a few minor observations of my own soon.

First, a basic but intriguing observation from Mathoverflow user “dorky”: The polynomials a, b and c are homogeneous with respect to the grading where \deg(x) = -1, \deg(y) = 1 and \deg(z)=2; their degrees are \deg(a) = 2, \deg(b) = 1 and \deg(c) = -1. I’m not sure what to make of this, but it surely matters.


Some computations by me: If you eliminate any two of the variables (x,y,z), you get a cubic relation in the remaining variable. Here they are

−2c+(4−3bc)x+(16a−b2−18abc+b3c+27a2c2)x3(−18ab+b3+27a2c)+18ay−3by2+2y3(really long)+8z3\begin{matrix} -2 c+(4 – 3 b c) x + (16 a – b^2 – 18 a b c + b^3 c + 27 a^2 c^2) x^3 \\ (-18 a b + b^3 + 27 a^2 c)+18 ay-3 b y^2+ 2y^3 \\ (\text{really long}) + 8 z^3 \\ \end{matrix}

I’m leaving out the “really long”, because it is really long and I suspect we don’t care about the details. Put

Δ=16a−b2−18abc+b3c+27a2c2\Delta= 16 a – b^2 – 18 a b c + b^3 c + 27 a^2 c^2 ,

the leading coefficient of the x cubic. Then the discriminants of the three cubics are \Delta p^2, \Delta q^2, \Delta r^2 where

pamp;=amp;8−9bc+27ac2qamp;=amp;bramp;=amp;(really long)\begin{align*} p &amp;=&amp; 8 – 9 b c + 27 a c^2 \\ q &amp;=&amp; b \\ r &amp;=&amp; (\text{really long}) \\ \end{align*}

The polynomials (p,q,r) have no common zeroes. Roughly speaking, our map should have special behavior over the loci \Delta=0, p=0, q=0 and r=0. The fact that $p$, $q$ and $r$ each appear cubed means that the variables x, y and z should have three fold branching over the loci p=0, q=0 and r=0 (respectively).

I’m having trouble visualizing what happens over \Delta=0 — since the leading coefficient of the x cubic drops out, the map is 2 to 1 rather than 3 to 1 over this point. But, at the same time, the y and z cubics have a multiple root at the points of \Delta=0. Does anyone see how to visualize this?

Any other insights?