Disclaimer: due to current events, I have not been able to devote as much time to lecture notes preparation as I would have liked, so I apologize in advance for the unpolished nature of the text below, which has been largely recycled from previous lecture notes I have written.
This is the first set of lecture notes for my graduate course 247A, “Fourier analysis”. The course name is rather general, but I will focus the course not on the Fourier transform per se, but on the closely related topic of real variable harmonic analysis, with a particular emphasis on Calderón–Zygmund theory, which underlies basic tools in PDE such as the theory of Sobolev spaces.
To avoid confusion at the outset, let us make the distinction between real-variable harmonic analysis and abstract harmonic analysis, which are only distantly related to each other despite the similar names. Abstract harmonic analysis, roughly speaking, is the extension of the classical theory of the Fourier transform to other domains, such as locally compact abelian (LCA) groups, non-abelian Lie groups, or symmetric spaces, and typically involves a blend of representation theory, group theory, and analysis. Real-variable harmonic analysis, by contrast, tends to work on classical domains, such as a Euclidean space
, a torus
, or a lattice
, although many of the techniques can extend to more general domains (e.g., to Riemannian manifolds). While the Fourier transform often plays a prominent role (in particular, by setting the stage for time-frequency analysis and enabling various decompositions or other transforms that involve frequency space or phase space in addition to physical space), real-variable harmonic analysis is often focused on estimating other transforms or expressions that often interact well with the Fourier transform, but need not explicitly invoke it. Examples include the Hilbert transform

(where we have made the somewhat arbitrary decision to omit the normalizing constant

) or the
Hardy-Littlewood maximal function 
These operators do not at first glance seem related to the Fourier transform, but as we shall see in later notes, the maximal function can be used to control the size of the Hilbert transform, and the Hilbert transform is a Fourier multiplier, with

for any reasonable function

, where we use the normalization

.
A typical question in harmonic analysis is the following: let
be some function on a standard domain (such as Euclidean space), and let
be an explicit transform of
(e.g., the Hilbert transform
or maximal function
). To what extent is the “size” of
controlled by the “size” of
? The value of such bounds often lies in the general nature of the input function
; some mild regularity or decay hypotheses might be imposed on
, but beyond that the function
is typically not required to have a very structured form (in particular, it need not be describable by any closed-form expression).
In many situations the transform
being studied is linear or sublinear, in which case the natural type of bound to ask is a linear bound

for suitable function space norms

(e.g.,

norms), and

is some bound. Depending on the application, we may be interested in various levels of precision regarding the bound

:
- (a) Optimal bounds, in which we seek the exact optimal value of
(i.e., the operator norm of
). For instance, the optimal
constant for the Hilbert transform
is exactly
, and the optimal
constant for
in one dimension turns out to be
(a result of Melas). - (b) Bounds accurate up to absolute constants (or maybe constants that can depend on basic parameters such as the ambient dimension).
- (c) Bounds in which we are willing to accept “logarithmic type losses” such as
or
in auxiliary parameters, such as a scale parameter
.
All three regimes are interesting, but we will focus in this class on the regime (b), where we can “afford” to lose absolute constants in the bounds, but will work hard to avoid any logarithmic losses. In particular, significant effort will be devoted in this class to avoiding “logarithmic pileups of scales”, in which the contributions of different dyadic scales such as
for
all potentially contribute an equal amount that “interfere constructively” to cause a logarithmic divergence. This can be unnecessarily conservative when one is in regime (c) (which is for instance the situation in modern topics such as restriction theory or the Kakeya conjecture); nevertheless, the general skills gained by trying to not lose even a logarithmic factor in the bounds are often valuable in these other types of analysis.
When dealing with linear or sublinear problems, it is natural to try to decompose the initial function
into various smaller components
by some decomposition
, so that the transformed function
can be controlled by more tractable expressions
in various ways (e.g., via the triangle inequality, by Bessel type inequalities, or by the more modern technique of decoupling inequalities). In short, the subject tends to proceed by a divide and conquer philosophy: it is generally preferable to replace a simple-looking but hard-to-estimate expression with a large, messy-looking combination of expressions that are easier to estimate. As such, the aesthetics of the subject are almost the reverse of those in the more algebraic portions of mathematics, in which progress is often made by making the expressions involved look as simple and unified as possible.
One of the main themes in this classical type of harmonic analysis is the struggle to understand the effect of two phenomena in integrals or sums: singularity and oscillation. The Hilbert transform (1) is a quintessential example of a singular integral, which combines both features: the non-locally integrable nature of the kernel
provides the singularity, but the sign change from
to
provides the oscillation. Classically, the interplay between these two phenomena can be tamed by analyzing the behavior of this operator both in the time (or “physical”) domain and in frequency (or “Fourier”) domain; in particular, the fact that singular integral operators such as the Hilbert transform are simultaneously a well-behaved Fourier multiplier and is “pseudo-local” in physical space lies at the heart of the standard Calderón–Zygmund theory for such operators. This dovetails nicely with more modern “time-frequency analysis” approaches to the subject, which can also handle other interesting operators, such as restriction or Bochner–Riesz operators via tools such as the wave packet decomposition, although these will be outside the scope of this course.
In this initial set of notes I will ignore the effect of oscillation, and develop some tools, such as interpolation theory, which can help control non-oscillatory sums and integrals if they are not too singular. Here, the focus will be on rearrangement-invariant spaces, such as the Lebesgue spaces
and their weak variants
, which are function spaces that are useful for measuring how “singular” or “decaying” various functions are, but do not pay attention to how they oscillate or where their mass is distributed. As such, these spaces do not capture the underlying geometry of the domain, which also plays an essential role in the subject; but it is nevertheless essential to have a good base understanding of the rearrangement-invariant theory before moving on to the more delicate aspects of harmonic analysis that are sensitive to rearrangements.
We will use the following asymptotic notation throughout the course:
,
, or
denotes the assertion that
for some constant
, and write
for
. If we permit this constant
to depend on some ambient parameters, we indicate this by subscripts; for instance,
or
denotes a bound of the form
for some constant
that can depend on
and
. As indicated above, in this course we will generally not dwell much on exactly what these constants
are, or attempt to optimize them.
— 1.
norms —
Suppose one has some measurable function
on some measure space
. (Here we will follow the common practice if identifying functions that agree almost everywhere; in particular, we will be content to work with functions that are undefined on a set of measure zero. Also, while we work here with complex-valued functions throughout, most of the discussion here is also valid for real-valued or vector-valued functions.) Informally speaking, to measure how “big” such a function is, there are two (imprecisely defined) basic statistics to be aware of:
- The height or amplitude of the function, which describes what the typical size of the magnitude
is for
in the “dominant” component of the support of
; and - The width of the function, which describes the measure of this dominant component.
Example 1 (Informal) Given a Gaussian wave packet type function 
on
for some
and
, this function has magnitude
on the ball
, which has volume
(if we allow constants in the informal
notation to depend on the dimension
), so such a function has height
and width
.
Example 2 (Informal) The function
on
, which is implicitly involved in the definition of the Hilbert transform (1), does not have a clear amplitude or width as is. However, if one performs a dyadic decomposition 
where we use
to denote the indicator of a statement
(equal to
when
is true and
otherwise), then each component 
of this decomposition has height
and width
. Thus, while this function can be viewed as a superposition of components of various heights and widths, rather than a single such component.
These informal concepts of height and width are too imprecise to work with in practice. Experience has shown that a convenient proxy for these concepts are the
norms of a function
, defined for
as

and for

as

where

denotes the
essential supremum of the function

with respect to the measure

. Often we abbreviate

as

,

,

, or just

(and abbreviate

as

) when the missing arguments are clear from context. (For instance, when working with Euclidean spaces

, the measure

is understood to be Lebesgue measure, and

the Lebesgue

-algebra, unless otherwise specified.) In terms of the width

and height

of a function

, one heuristically has

for both finite and infinite values of

, with the convention that

is equal to

when

is positive and

when

is zero. In the case of a step function

(where now

is the indicator function of a measurable set

), this heuristic becomes exact:

The function space
is defined as the set of all measurable functions
for which the
norm is finite, up to almost everywhere equivalence, though we will often abuse notation by identifying a function with its almost everywhere equivalence class.
In the case where
is discrete and
is counting measure, we abbreviate
as
, or even just
.
Example 3 Let
. On a Euclidean space
, the function
lies in
(with a norm of
) if and only if
, while the function
lies in
(with a norm of
) if and only if
. The function
does not lie in any
, although it only fails “logarithmically” to lie in
. Thus we see that control in
for high
rules out severe local singularities at a point, while control in
for low
rules out insufficiently rapid decay at infinity.
As is well known (see these previous notes) these function spaces enjoy many useful properties:
Theorem 4 (Basic properties of
spaces)
Remark 5 Closely related to (iii) is the fact that the dual of
can be identified with
when
(with the additional hypothesis that
is
-finite if
), but in practice the relation (4) will already be good enough for our purposes.
Remark 6 The Banach space property gives us the basic triangle inequality 
for both finite and infinite collections of functions
when
, where in the infinite case the assertion is that if the right-hand side is finite, then the series
is absolutely convergent almost everywhere, and obeys the above inequality (so in particular is in
). For
, this inequality fails (can you come up with a counterexample?), but one has the weaker
-triangle inequality 
in this case, which follows easily from iterating the easy observation that
for any complex numbers
, which in turn ultimately stems from the complex triangle inequality
and concave nature of
for
. In particular, for a finite sum
, another application of Hölder’s inequality gives the quasi-triangle inequality 
for
, which is not too much worse than (5) when
is not too large.
Exercise 7 Give an example to show that the quantity
in (7) cannot be replaced by any smaller quantity.
Exercise 8 For
a simple function, verify that
, and that
, where
. For this reason, the measure of the support of
is sometimes referred to as the
norm of
, though it would be more accurate (though confusing) to refer to it as the
power of the
norm.
Remark 9 Note that Hölder’s inequality is not just symmetric under the homogeneities
and
of the functions, but also under the homogeneity
of the underlying measure. This latter symmetry demonstrates why the condition
is necessary. (The first two symmetries demonstrate why
appears the same number of times on both sides of the inequality, and similarly for
.)
In the case of Euclidean space, the measure homogeneity symmetry
is equivalent to the scaling symmetry
for
, as the Jacobian of this map is
. But the point is that by manipulating the measure directly, one still enjoys this symmetry even when no scaling operation is present.
It is instructive to try to understand inequalities such as (3) using the height-width heuristic introduced previously. Suppose informally that
have heights
,
,
and widths
,
,
respectively. Then one expects the heights to be related by the formula

What about the widths? Heuristically, the region that

concentrates in ought to be a subset of the region that

concentrates in, so

and similarly with

replaced by

. We can combine these bounds as

The bound
(3) then is morally

which on applying the previous bounds and
(2) should simplify to

But this is clear by bounding

by

for the first factor on the left-hand side, and by

for the second factor. Thus we see that the key geometric input that is morally driving the Hölder inequality is the simple fact that the concentration region of the product is contained in the concentration regions of the factors.
Exercise 10 Determine the cases for which (3) holds with equality (dealing with edge cases such as when one or more of
equal infinity as appropriate). Discuss how your conclusions align with the heuristic analysis presented above.
Exercise 11 If
, determine the cases for which (5) holds with equality. What changes when
or
?
Exercise 12 Show that Hölder’s inequality is equivalent to the log-convexity of
norms: 
(For technical reasons one needs to first reduce to the case where
has finite measure, and then
and
are everywhere non-vanishing simple functions. Now consider the convexity of
with respect to a measure
for some suitable exponents
.)
Exercise 13 (Direct approach to log convexity) Differentiate
twice with respect to
and show that this is non-negative (take
to be a non-zero simple function with finite measure support to avoid technicalities). This is an example of a monotonicity formula method — deriving estimates from a monotonicity property, which in turn follows from the non-negativity of a derivative.
You will see that this approach is surprisingly messy. For a much slicker proof, observe that (8) enjoys homogeneity symmetry in both
and
, which lets one normalise both
and
to equal one. Thus the task is now to show that if
, then
for all
between
and
. This can be done by the pointwise convexity of
, or more precisely the estimate

the observant reader will note that this is merely the proof of Hölder’s inequality in disguise.
Let us now give a more unusual proof of the log-convexity which does not appeal to any pointwise convexity estimate, instead combining the “divide and conquer” strategy with an elegant (and rather cheeky) “tensor power trick“. Again normalise
. We split
into a broad flat piece and a narrow tall piece

which are disjoint, and thus

What we are doing here is exploiting some very basic intuition about

norms, namely that

bounds for large

tend to exclude tall narrow spikes, whereas

bounds for small

tend to exclude short broad tails. Of course, either sort of bound would exclude tall broad functions, and neither excludes narrow short functions. Once again, this intuition can be buttressed by considering the special case of step functions.
When
, then
, and when
, then
. Thus we end up with

The above argument (which is a prototype of the
real interpolation method) obtained an estimate which is off by a factor of two from what we wanted; this is a typical feature of the method. However we can recover this factor for free by the following
tensor power trick. Let

be a large integer. We replace the measure space

by its

power

using the product measure construction, and similarly replace

with its tensor power

, defined by

One then observes that

Now we apply the preceding arguments to

instead of

to deduce that

which on taking

roots gives

Now the left-hand side is independent of

; take limits as

and we obtain

as desired.
The tensor power trick can be viewed as another application of symmetry: if an estimate is invariant under raising to a tensor power, then one can automatically replace all absolute constants with
; thus we obtain the “free lunch” of deducing a bound with an explicit constant
, from a bound with an unspecified constant (or even with “logarithmic losses”). Contrapositively, if an estimate is invariant under tensor power, then a weak counterexample (which shows that the constant must exceed one) can be amplified into a strong counterexample (which shows that no finite constant suffices) by tensor powering. The tensor power trick seems like a magical trick at present, but is actually exploiting some basic results in information theory such as the Shannon entropy inequalities and the central limit theorem; it also combines well with virtually any inequality which involves Gaussians. Unfortunately due to lack of time we will not be discussing these beautiful topics further in this course. At any rate one sees the power of abstraction in this tensor power trick. (One could similarly perform this trick in
, so long as the constants only grew sub-exponentially in the dimension
.)
The final proof of log-convexity of the norm that we give here proceeds via complex analysis, and the maximum principle — which in many ways is a complex analogue of convexity (or subharmonicity). We need the following result from complex analysis, namely a form of the Phragmén–Lindelöf principle.
Lemma 14 (Three lines lemma) Let
be a complex-analytic function on the strip
, which is of at most double-exponential growth, or more precisely
for some
. Suppose that we have the bounds
when
and
when
. Then we have
for all
in the strip.
Remark 15 The rather strange sub-double-exponential hypothesis here is completely sharp, as the example
shows. Note in this hypothesis that we allow the implied constants in the asymptotic notation to depend on
, but the hypothesis is qualitative rather than quantitative: the value of these constants is irrelevant for the final conclusion, as long as they are finite. In practice, these sorts of qualitative hypothesis are usually easy to establish (especially when compared to quantitative estimates) by restricting, smoothing, or damping to a nice class of functions, or by smoothing out or discretising various operators and domains. See for instance the proof of this very lemma in which we upgrade “for free” a weak qualitative bound (sub-double-exponential growth) to a strong qualitative bound (decay at infinity).
Proof: The hypotheses and conclusion of the lemma are invariant under the operation of multiplying
by a constant (and adjusting
appropriately). So we may normalise
. Similarly, the hypotheses and conclusion of the lemma are invariant under the operation of multiplying
by an exponential
for some real
. Using this, one can also normalise
. So now
is bounded by
on both sides of the strip and we want to show it is bounded by
inside the strip.
Let us first assume that
is much better than exponential growth, namely that it goes to zero at infinity. Then for all sufficiently large rectangles
the complex-analytic function
is bounded by
on all four sides of this rectangle, and hence in the interior also by the maximum principle, and we are done by setting
.
Now let us handle the general case; as is usual when removing a qualitative assumption, we do this by a limiting argument. We replace
by
; a little complex arithmetic shows that this converts the almost double-exponentially growing function
to one which is still complex analytic but is now decaying at infinity. It is still bounded by
at both sides of the strip, and hence by
in the interior also by the previous argument. Now take
to conclude the claim.
Exercise 16 Suppose that
is analytic on the strip
, obeys the sub-double-exponential bound
on the strip, and obeys the polynomial bounds
on the sides of the strip. Show that it obeys the polynomial bound
on the interior of the strip also.
To apply the three-lines lemma to prove (8), take
to be a simple function (with finite measure support) and consider the entire function

This function has exponential growth at most (because of the qualitative assumption that

is simple with finite measure support), and is bounded by

on the lines

and

, and hence (by a trivially rescaled version of the three lines lemma) bounded by

on the strip inside the lines. In particular it is bounded by

at

, which gives the claim for simple functions. The claim for more general functions (dropping the qualitative assumption of simpleness and finite measure support) then follows by a standard limiting argument (using for instance monotone convergence) which we leave as an exercise.
The above argument is a prototype of the complex interpolation method. As one can see, it can give slightly sharper results than the real interpolation method (though using the tensor power trick the real method can sometimes “catch up”), but on the other hand requires the quantities being studied to depend complex-analytically on a parameter rather than (say) real-analytically.
Having conclusively demonstrated the log-convexity (8) in multiple ways, let us now give some quick applications. It shows that control on two extreme
norms implies control of the intermediate
norms. Under additional assumptions on the measure space
, one of these extremes is not necessary. If the measure space is finite in the sense that
(thus prohibiting functions from being arbitrarily broad), then higher
norms control lower ones:

Indeed this is trivial when

, and the general case then follows by convexity. The bound
(9) can also be usefully written in terms of averages: if we write

for

, then we see that higher

averages control lower

averages:

One way to view this is that as one lowers the exponent

, the exceptionally large values of

become less important, leaving the small values of

to dominate. By restricting

to its support

one can refine
(9) to

(Note that this is a limiting case of log-convexity at the exponent

, in view of Exercise
8). In the converse direction, if the measure space is
granular in the sense that one has a lower bound

for all sets

of positive measure, then functions are prohibited from being arbitrarily narrow, and lower norms control higher norms:

This can be seen by first checking the

case, and then using log-convexity to get the remaining cases. In particular, in the

spaces we see that

for

. For

spaces on

points, we thus have (non-matching) upper and lower bounds


and

norms are comparable to some extent, but the comparability gets worse as

or as

and

get further apart.
Exercise 17 Heuristically justify the bounds (9), (10) by appealing to the informal notions of width and height of a function.
Exercise 18 When does equality occur for either of the inequalities in (10)? Note how the example that attains the lower bound is in many ways the “opposite extreme” to the example which attains the upper bound.
Lebesgue measure on Euclidean spaces
with the usual Borel or Lebesgue
-algebra is not granular. However one can create granularity by coarsening the
-algebra. For instance, if we let
be the
-algebra generated by the lattice unit cubes
for
, then we have granularity with constant
, and now lower
norms of functions
control higher ones — but only for functions which are measurable with respect to this algebra, i.e. only for functions which are constant on each lattice unit cube. (This is the first time we have actually manipulated the
-algebra
to say something non-trivial, as opposed to manipulating
,
, or
.) Thus we see that local constancy of functions can lead to additional estimates on
norms. Later on we shall see that frequency localisation achieves a similar effect as local constancy, as quantified by Bernstein’s inequality; this is a concrete manifestation of the famous Heisenberg uncertainty principle.
Finiteness and granularity of the measure space prevent a function from being too broad or too narrow respectively. Similar things happen when a function is being prevented from being too tall or too short; for instance if
is bounded above by a constant
, then we have

(this is just log-convexity at the

exponent), while if

is bounded below by

on its support, then we have the reverse inequality

There are two obvious algebraic identities involving
norms which are worth knowing. The first is that one can interchange
sums with
integrals for any
, in the sense that

this is just an application of the Fubini-Tonelli theorem. Secondly, exponents can pass through

norms by changing the exponent: for any

we have

We shall use both of these identities in the sequel without further comment.
Exercise 19 Establish the bound 
for any measurable
and any
.
— 2. Lorentz spaces —
Recall that the weak
norm
of a function
is defined for
as

Since

for any

and

, we obtain
Chebyshev’s inequality 
(the case

is also known as
Markov’s inequality). When

we adopt the convention that

.
We define weak
or
to be the space of all functions with finite
norm, with the usual abbreviations. We sometimes refer to
as strong
to distinguish it from weak
.
Example 20 On a Euclidean space
, the power function
lies in weak
if and only if
. Indeed one can think of a weak
function as a function which is pointwise dominated in magnitude by a rearrangement of (a multiple of)
.
Suppose
. From elementary calculus we have

and hence on integration and Fubini’s theorem

To summarise, we have

and

for

. These two identities motivate introducing the
Lorentz (quasi-)norm 
for

and

by

Thus for instance

norm is identical to the

norm. We shall abbreviate

by

,

, or even

when there is no chance of confusion.
Remark 21 For various reasons it is not worth trying to define Lorentz norms when
, although we will use the convention
. The most important values of
, in descending order, are
,
,
, and
; the other cases essentially never occur in applications.
Remark 22 The factor
is inconsequential, but is traditional in order to maintain compatibility with the strong
norm. But in practice the exact form of the Lorentz norm is not important; there are many formulations which are equivalent up to constants, and one generally just picks the formulation which is most convenient. The measure
is of course multiplicative Haar measure on
. One can interpret the equivalence of (i)-(iii) below by making the change of variables
, so that the Haar measure just becomes Lebesgue measure in
(modulo an inessential constant) and then passing from continuous
to discrete
.
Exercise 23 If
is a monotone non-increasing function, show that 
(Depending on how you prove this, it may be convenient to first prove this for smoother
, such as diffeomorphisms with strictly negative derivative, in order to apply an inverse function theorem.)
Exercise 24 Show that a step function of height
and width
has an
norm of
for any
and
.
Exercise 25 For any
and
show that 
It is obvious that these norms are both rearrangement-invariant and monotone. To get a better intuitive handle on what the
norm represents, we need some more definitions.
Remark 26 It is a little dangerous to put fuzzy notation such as
within a definition; if multiple quasi-step functions appear in an argument, the question then arises as to whether the implied constants are uniform. In our applications, the implied constants here are true absolute constants (like
and
) so this will not be an issue.
Remark 27 From the binary expansion of the unit interval
we see that a non-negative sub-step function
of height
and width
can always be decomposed as
where
is an actual step function of height
and width at most
. By homogeneity we have a similar statement for other heights. Because of this, bounds on step functions tend to automatically extend to sub-step functions (and hence quasi-step functions) without difficulty.
Just like actual step functions, the
norm of sub-step and quasi-step functions are well controlled; a sub-step function of height
and width
has
norm
, while a quasi-step function has
norm
– almost exactly like actual step functions. In the converse direction, it turns out that every
function can be decomposed as an
sum of “very different”
-normalised sub-step or quasi-step functions.
Theorem 28 (Characterisation of
) Let
be a function, let
and
, and let
. Then the following are equivalent up to changes in the implied constants:
Remark 29 The formulations (ii), (iv) are useful when trying to use an
bound on
; the formulations (iii), (v) are useful when trying to obtain an
bound on
. Heuristically, the above theorem is trying to say the following. If
is a quasi-step function of height
and width
, then
. But if
is instead the sum
of quasi-step functions of height
and width
, and either the heights or the widths are sufficiently variable in
(e.g. one or the other grows like a power of two), then
.
Proof: We may use homogeneity symmetry to normalise
. The implications
and
are trivial. To see that (i) implies (ii), set
and
. (This is the “vertically dyadic layer cake decomposition”.) The only thing that requires nontrivial verification is (12); but one easily verifies that
![\displaystyle 2^m W_m^{1/p} \lesssim_{p,q} \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^q([2^{m-2}, 2^{m-1}],\frac{d\lambda}{\lambda})}](https://s0.wp.com/latex.php?latex=%5Cdisplaystyle++2%5Em+W_m%5E%7B1%2Fp%7D+%5Clesssim_%7Bp%2Cq%7D+%5C%7C+%5Clambda+%5Cmu%28+%5C%7B+%7Cf%7C+%5Cgeq+%5Clambda+%5C%7D+%29%5E%7B1%2Fp%7D+%5C%7C_%7BL%5Eq%28%5B2%5E%7Bm-2%7D%2C+2%5E%7Bm-1%7D%5D%2C%5Cfrac%7Bd%5Clambda%7D%7B%5Clambda%7D%29%7D&bg=ffffff&fg=000000&s=0&c=20201002)
and the claim follows by summing this in

.
Similarly, to see that (i) implies (iv), define

note that this is a non-increasing function of

, which goes to zero as

(this comes from the hypothesis that

is finite). We then define

(This is the “horizontally dyadic layer cake decomposition”.) The only non-trivial thing to verify is
(14). But one easily verifies the telescoping estimate
![\displaystyle \begin{array}{rl} H_n 2^{n/p} &= (H_n^q 2^{nq/p})^{1/q} \\ &= (\sum_{k=0}^\infty (H_{n+k}^q - H_{n+k+1}^q) 2^{nq/p} )^{1/q} \\ &\lesssim_{p,q} (\sum_{k=0}^\infty 2^{-kq/p} \| \lambda 2^{(n+k)/p} \|_{L^q([H_{n+k+1}, H_{n+k}],\frac{d\lambda}{\lambda})}^q)^{1/q}\\ &\lesssim_{p,q} (\sum_{k=0}^\infty 2^{-kq/p} \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^q([H_{n+k+1}, H_{n+k}],\frac{d\lambda}{\lambda})}^q)^{1/q} \end{array}](https://s0.wp.com/latex.php?latex=%5Cdisplaystyle++%5Cbegin%7Barray%7D%7Brl%7D++H_n+2%5E%7Bn%2Fp%7D+%26%3D+%28H_n%5Eq+2%5E%7Bnq%2Fp%7D%29%5E%7B1%2Fq%7D+%5C%5C+%26%3D+%28%5Csum_%7Bk%3D0%7D%5E%5Cinfty+%28H_%7Bn%2Bk%7D%5Eq+-+H_%7Bn%2Bk%2B1%7D%5Eq%29+2%5E%7Bnq%2Fp%7D+%29%5E%7B1%2Fq%7D+%5C%5C+%26%5Clesssim_%7Bp%2Cq%7D+%28%5Csum_%7Bk%3D0%7D%5E%5Cinfty+2%5E%7B-kq%2Fp%7D+%5C%7C+%5Clambda+2%5E%7B%28n%2Bk%29%2Fp%7D+%5C%7C_%7BL%5Eq%28%5BH_%7Bn%2Bk%2B1%7D%2C+H_%7Bn%2Bk%7D%5D%2C%5Cfrac%7Bd%5Clambda%7D%7B%5Clambda%7D%29%7D%5Eq%29%5E%7B1%2Fq%7D%5C%5C+%26%5Clesssim_%7Bp%2Cq%7D+%28%5Csum_%7Bk%3D0%7D%5E%5Cinfty+2%5E%7B-kq%2Fp%7D+%5C%7C+%5Clambda+%5Cmu%28+%5C%7B+%7Cf%7C+%5Cgeq+%5Clambda+%5C%7D+%29%5E%7B1%2Fp%7D+%5C%7C_%7BL%5Eq%28%5BH_%7Bn%2Bk%2B1%7D%2C+H_%7Bn%2Bk%7D%5D%2C%5Cfrac%7Bd%5Clambda%7D%7B%5Clambda%7D%29%7D%5Eq%29%5E%7B1%2Fq%7D+%5Cend%7Barray%7D+&bg=ffffff&fg=000000&s=0&c=20201002)
and the claim follows by summing this in

and interchanging the summation signs. (We leave to the reader how to modify the above argument to handle the case

.)
It remains to show that (iii) implies (i) and (iv) implies (i). Suppose first that (iii) holds. It is clear that for any
we have

and hence
![\displaystyle \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^q((2^m,2^{m+1}],\frac{d\lambda}{\lambda})} \lesssim_{p,q} 2^m (\sum_{k=0}^\infty \mu(E_{m+k}))^{1/p}](https://s0.wp.com/latex.php?latex=%5Cdisplaystyle++%5C%7C+%5Clambda+%5Cmu%28+%5C%7B+%7Cf%7C+%5Cgeq+%5Clambda+%5C%7D+%29%5E%7B1%2Fp%7D+%5C%7C_%7BL%5Eq%28%282%5Em%2C2%5E%7Bm%2B1%7D%5D%2C%5Cfrac%7Bd%5Clambda%7D%7B%5Clambda%7D%29%7D+%5Clesssim_%7Bp%2Cq%7D+2%5Em+%28%5Csum_%7Bk%3D0%7D%5E%5Cinfty+%5Cmu%28E_%7Bm%2Bk%7D%29%29%5E%7B1%2Fp%7D&bg=ffffff&fg=000000&s=0&c=20201002)
and so on taking

summation in

it would suffice to show that

Raising to the

power we rewrite as

But from the hypothesis we have

and hence on shifting

by

The claim then follows from
(5) or
(7).
Now suppose that (iv) holds. Observe that for any
we have

where

are the modified heights

Indeed, if

for some

then one easily verifies that

and hence

. The shifting trick and triangle inequality argument used previously shows that

obeys the same bound
(14) as

, thus

We now compute

as desired. (We leave to the reader how to modify the above argument to handle the case

.)
Remark 30 For future reference we make the technical remark that if
is a simple function, then in the horizontal and vertical decompositions in the above theorem, only finitely many of the
are non-zero.
Remark 31 Suppose that the ratio between the tallest height and lowest non-zero height of a function
is
(i.e. there exists
such that
whenever
is non-zero). Then the above theorem shows that two different Lorentz norms
,
with the same primary exponent
only differ by multiplicative powers of
. Similarly if the broadest width and narrowest width of a function differs by
(e.g. if
is equal to
times the granularity
of
). What this indicates is that the secondary exponent
in the Lorentz norms only offers “logarithmic correction” to the Lebesgue norms
; in contrast, (10) shows that varying the primary exponent
leads to polynomial-strength changes in the norm. So as a first approximation (ignoring logarithms) one can pretend that
. Note also that for quasi-step functions, the
norms barely depend on
at all.
One easy corollary of the above theorem is that the
quasi-norm is indeed a quasi-norm, and in particular that

for any

; this can be seen for instance by using the equivalence of (i) and (iii). Another easy consequence is that the simple functions are dense in

.
Exercise 32
Exercise 33 For each integer
, let
be a quasi-step function of height
and width
for some
. Show that 
for all
. If instead
is a quasi-step function of height
and width
for some
, show that 
for all
. What goes wrong when we remove the absolute values on the
? (This can be repaired by replacing the powers of
with powers of a sufficiently large constant (depending on the implied constant in the definition of a quasi-step function) – why?)
A particularly useful consequence of the above theorem is a Hölder inequality for Lorentz spaces, due to O’Neil.
Theorem 34 (Hölder’s inequality in Lorentz spaces) If
and
obey
and
then 
whenever the right-hand side norms are finite.
Proof: We may normalise
, and drop the dependence of the implied constants on
for brevity. By the equivalence of (i) and (v) in Theorem 28 we may dominate
and
where
and

Then we have

By the quasi-triangle inequality and monotonicity it suffices to show that

and

By symmetry it suffices to consider the

component. Here we observe that

has measure at most

, so by the equivalence of (i) and (v) in Theorem
28 
But by the ordinary Hölder inequality

shifting the second

by

we conclude

The claim now follows from Exercise
32.
One corollary of this Hölder inequality is that
functions are absolutely integrable on sets of finite measure whenever
.
Now we consider dual formulations of the
norms. The case
is fairly straightforward:
Exercise 35 (Dual formulation of weak
) Let
. Then for every
in
, we have 

Also show that the hypothesis
can be dropped if one instead assumes
to be non-negative.
The right-hand side of (16) is clearly a semi-norm at least on
. This leads in particular to a quasi-triangle inequality

for any

.
Remark 36 It is worth comparing (16) to (4). In (4), one takes the inner product of
against all
-normalised functions, and the worst inner product becomes the
norm. In (16), one only takes the inner product of
against the
-normalised step functions
. This is fully consistent with the fact that the
norm is stronger than the weak
norm.
Exercise 35 can be rephrased as follows: if
for some
and
, then the following two statements are equivalent (up to changes in the implied constants):
-
. -
for all sets
of finite measure.
Unfortunately this equivalence breaks down at
or below (consider for instance the weak
function
on
, which is not even locally integrable when
). However, one does have a substitute:
Exercise 37 Let
,
, and
. Show that the following are equivalent (up to changes in the implied constant):
Hint: It may be instructive to work out the example
on
by hand to get a sense of what is going on; this should suggest how to prove things in general. The proof is slightly simpler in the case when
is non-negative, so you may want to try that case first. Comment on how this result implies Exercise 35 (or its equivalent version discussed shortly afterwards) when
.
Exercise 38 Let
be functions with
. Show that 
thus the weak
quasinorm only fails to be a norm “by a logarithm”. Show with an example that the
cannot be removed.
For more general
spaces, we have
Theorem 39 (Dual characterisation of
) Let
and
. Then for any
, 
Again, the hypothesis
can be dropped if
is non-negative and
is restricted to also be non-negative.
Thus the
quasi-norm is in fact equivalent to a norm when
and
. In particular, weak
is equivalent to a normed space when
. (For
, weak
fails to be normable “by a logarithm”; see Q3.) As with other dual characterisations, one can restrict
to a dense subclass of
, for instance simple functions with finite measure support.
Proof: To obtain the
part of this theorem, we simply estimate

and use Theorem
34. To obtain the

part, we normalise

. It then suffices by homogeneity to find

with

and

.
The case
follows from (16), so let us take
. We will just give the proof in the case
; the case
is trickier, and a partial argument is given in the exercises. By the equivalence of (i) and (ii) in Theorem 28 may write
where
is a quasi-step function of height
and width
with disjoint supports such that the sequence
has an
norm of
. Now take

where

adopting the obvious convention that

when

. Then (because of the disjoint supports)

But since

has height

and width

,

and so

To conclude it will suffice to show that

If

is the support of

, then we have the pointwise bound

and the measure bound

.
At this point we would like to apply Theorem 28, but neither the height nor width of
is necessarily a power of
. But we can remedy this by introducing the modified heights

We have

, and so the

increase geometrically. It then suffices to show that

By refining the

by a constant factor we can make each

at least twice as large as the previous, and so by applying the equivalences of (i) and (iii) in Theorem
28 and the triangle inequality it suffices to show that

which we expand using our bound on

as

But from Hölder’s inequality (using the hypothesis

) and the

bound on

we have

summing this using the triangle inequality (and estimating the supremum by a sum) we obtain the claim.
The case when
are restricted to be non-negative can be deduced from the above result and a monotone convergence argument (representing
as a monotone limit of simple functions of finite measure support) which we leave as an exercise to the reader.
Exercise 40 A measure space
is said to be non-atomic if, for every measurable set
with
, there exists a measurable subset
such that
.
- (i) (Sierpinski’s theorem) Show that if
is non-atomic,
is measurable, and
, then there exists a measurable subset
of
such that
. (You may find it convenient to use Zorn’s lemma.) - (ii) Show that if
is non-atomic, then the decomposition in Theorem 28(iv) can be chosen so that for each
, either
vanishes, or
has support of measure
(not just bounded by
). - (iii) Establish Theorem 28 in the case that
and
is non-atomic, by using the modification of Theorem 28 indicated in the previous part of the exercise.
The duality can also be established for measure spaces that contain atoms, but this requires a more careful analysis that treats large atoms separately.
— 3. Orlicz spaces (Optional) —
So far we have studied the Lebesgue spaces
, together with the more general Lorentz spaces
, which includes weak
as a special case. These spaces are all rearrangement-invariant and monotone. There is a different generalisation of the Lebesgue spaces
, the Orlicz spaces
, which are also rearrangement-invariant and monotone, and which are occasionally useful. (There is a common generalisation of both, the Lorentz-Orlicz spaces, but these occur very rarely in applications.)
The motivation for Orlicz spaces starts with the trivial observation that if
, then

Inspired by this, we generalise by letting

be a function (with some additional properties to be selected shortly) and ask if we can find a norm

which obeys the property

Since norms need to be homogeneous, this would imply

for all

. In particular, if

, then we need

To ensure this property it is thus very natural to require that

be
increasing. Also to deal with the zero norm case one typically requires

.
Next, in order for
to be a norm, the unit ball
needs to be convex. Looking at (17), we see that this will indeed be the case when
is itself convex. (Note that the proof of (5) was a special case of this argument).
We can put all the above discussion together and conclude: if
is increasing and convex with
, then the norm

is a norm on the space

.
As discussed above, the
spaces for
are examples of Orlicz spaces with
. The space
is not really an Orlicz space, but can be viewed the limiting case where
is infinite for
and zero for
(or more informally,
). Aside from the Lebesgue spaces, the most common Orlicz spaces which appear are
- The space
, defined as the Orlicz space with
; - The space
, defined as the Orlicz space with
; - The space
, defined as the Orlicz space with
.
The correction factors of
and
in the above functions should not be taken too seriously; note that if two functions
are comparable then their Orlicz norms are comparable also; a little more generally, if
, then
. It is the behaviour of
for large values of
which is the most important, although when
has infinite measure the behaviour at small values of
is also relevant.
Exercise 41 If
has finite measure, verify the relation 
which is another indication of the irrelevance of the low values of
in the finite measure case.
The final fact about Orlicz spaces that we give here is the duality relation. Suppose that
is increasing, convex, and is also superlinear in the sense that
. We can then define the Young dual
of
by the formula

the hypothesis that

is superlinear ensures that this function is well-defined. We may equivalently define

to be the smallest function for which one has the inequality

Exercise 42 If
, show that the Young dual of
is
. Show also that the Young dual of
takes the form
for
. What does the Young duals of
and
look like?
One can easily verify that
is also increasing, convex, and super-linear, and so the Orlicz norm
makes sense. From (18) and the triangle inequality it is immediate that

and hence by homogeneity we obtain the duality relation

whenever

and

.
Exercise 43 Establish the more precise relationship 
Exercise 44 Let
be superlinear, increasing, convex, and vanishing at the origin.
- (i) Show that the Young dual
of
is also superlinear, increasing, convex, and vanishing at the origin. - (ii) Show that
is the Young dual of
.
(It may help to view things geometrically, and in particular understanding
as parameterising the support lines of the graph of the convex function
.)
Exercise 45 When
has finite measure, show that the spaces
and
are dual to each other. What is the dual to
?
Exercise 46 Let
be a function on a measure space
of bounded measure
. Show that the following are equivalent (up to changes in the implied constants):
- (i)
. - (ii)
for all
. - (iii)
for all
.
Hint: You may find the Taylor expansion for
, together with the obvious bounds
for integer
(or Stirling’s formula, if you know what that is), to be useful.
Exercise 47 Obtain the analogue of Exercise 46 for the Orlicz space
.
— 4. Real interpolation —
So far we have only considered functions
on a single measure space
. Now we shall consider operators
which take functions on one measure space
to functions on another measure space
; the study of such operators is in fact a major focus of harmonic analysis. Ultimately we want to extend
to a standard normed vector space such as
, but in practice one has to initially first restrict attention to a dense subspace of functions, such as simple functions or test functions.
We are primarily interested in linear operators, thus
and
. But it is also worth considering the more general sublinear operators, in which

and we have the pointwise estimate

Apart from the linear operators, the next most important example of a sublinear operator is a
maximal operator 
where

are a collection (possibly countably or uncountably infinite, though in the latter case one has to take some care in ensuring measurability) of linear or sublinear operators. The third most important example is a
square function such as

More generally, one can consider a family

of operators indexed by some parameter

, and take

to be the norm in the

variable of

in some suitable norm. But the above three examples of linear operators, maximal operators, and square functions already cover the vast majority of applications.
Let
be exponents, and let
be sublinear. Let us define the following concepts.
Clearly, whenever
are fixed, strong-type implies weak-type and restricted strong-type, either of which imply restricted weak-type. In most applications, it is the strong-type bounds which are desired; however, we shall see in this section that the real interpolation method allows us to deduce strong-type bounds from weak-type, or even restricted weak-type bounds, as long as the strong-type bounds are an interpolant between the restricted weak-type bounds. This can be a very useful strategy, because (as we shall see in next week’s notes) weak-type or restricted weak-type bounds are easier to prove than strong-type estimate.
Let us first make a mild (and qualitative) assumption, namely that the form

is well-defined whenever

are simple functions with finite measure support. This is for instance the case if

is of restricted type

for some

and

, or restricted weak-type

for some

and

; thus in practice this assumption is easily satisfied. We observe that this form is non-negative, homogeneous and sublinear in both

and

:

The form

turns out to be a convenient way to understand the various types of

, and the ability to decompose both

and

independently is very useful in establishing interpolation type results. There is a near-symmetry between

and

here, if we could somehow take an “adjoint” of the operator

, but we will not explicitly exploit this symmetry here as it is not always available for sublinear operators (though the “duality” or “adjoint” trick is undoubtedly very powerful in the important linear case).
Let us look in particular at the form (23) applied to indicator functions, thus we look at the quantity
for
and
of finite measure. Now suppose that
had some strong type
bound for some
and
, say

for some

. Then in particular

and hence by Hölder’s inequality

Actually it is clear that strong type

is too much of an assumption; restricted type

would have sufficed for the conclusion. If

, we can relax things further to restricted weak-type:
Exercise 48 Let
,
, and
. Let
be a sublinear operator such that the form (23) is well-defined. Then the following are equivalent up to changes in the implied constant:
(Hint: use (16) and Remark 27.)
This already gives a simple version of real interpolation:
Corollary 49 (Baby real interpolation) Let
be a sublinear operator such that the form (23) is well-defined. Let
,
and
be such that
is restricted weak-type
with constant
(in the sense of (25)) for
. Then
is also restricted weak-type
with constant
for
, where 
and the implied constant depends on
.
Indeed, all we are using here is the obvious algebraic observation that if
and
, then
for all
.
The above corollary has two defects. Firstly, it can only conclude restricted weak-type rather than strong type. Secondly, the restriction
is inconvenient for many applications, since weak
bounds are actually rather common. To address the second concern, we have the following variant of Exercise 48:
Exercise 50 Let
,
, and
. Let
be a sublinear operator such that the form (23) is well-defined. Then the following are equivalent up to changes in the implied constant:
Hint: use Exercise 37.
Corollary 51 The hypothesis
in Corollary 49 can be weakened to
.
Now we can give a significantly more useful real interpolation theorem, which can interpolate between restricted weak-type estimates to obtain strong-type estimates.
Theorem 52 (Marcinkiewicz interpolation theorem) Let
be a sublinear operator such that the form (23) is well-defined. Let
,
and
be such that
is restricted weak-type
with constant
(in the sense of (25)) for
. Suppose also that
and
. Then for any
and
with
we have 
for all simple functions
with finite measure support, where
were defined in (26). In particular, if
, then
is strong-type
with constant
.
Proof: To simplify the notation let us suppress the dependence on
. Observe that the statement and conclusion of the theorem have several homogeneity symmetries. The most obvious one is that we can multiply
(and the
) by an arbitrary constant, but we may also multiply the measure
by a constant
(and the
by the constant
); similarly we may multiply
by
and
by
). Using these symmetries we may normalise
(and hence
for all
). Our task is then to show that

for all simple functions

with finite measure support.
Using Theorem 39 (and the hypothesis
), this is equivalent to showing that

for all simple functions

of finite measure support.
We currently have
and
. By using Corollary 51 to bring the restricted weak-type exponents
and
a bit closer to
we may assume that
as well. Applying Exercise 48 we conclude that

for all sets

of finite measure and

. From Remark
27 and sublinearity we conclude that

whenever

are sub-step functions of height and widths

and

respectively. We can of course pick the better of the two estimates, leading to

Now we can return to proving (27). By homogeneity we may normalise

We then apply Theorem
28 to decompose

,

with

sub-step functions of width

and heights

respectively, with the height bounds

where

are the sequences

Since

are simple functions, only finitely many of the

and

are non-zero by Remark
30. We now use sublinearity to estimate

and then use
(28) to obtain

We can write this in terms of

,

, and reduce to showing that

Because

and

, and because of the definitions of

, we can write the left-hand side as

for some non-zero

and some

depending only on

. We substitute

(thus

) and estimate this by

By
(29) and Hölder we see that the inner sum is

uniformly in

, and the claim follows.
Finally, if we specialise
and recall that the
norm will be dominated by the
norm for
, the last claim of the theorem follows.
There are many other variations on the real interpolation method, for instance an extension to multilinear operators, or to other function spaces. However, the basic method of proof is still the same: dualise, decompose all inputs, estimate each term as optimally as one can, and then sum.
One can illustrate the real interpolation method graphically using type diagrams. One plots all points
where the operator
is strong-type
or restricted weak-type
. Ignoring some of the technical hypotheses, the above interpolation theorems then essentially assert that the restricted weak-type diagram and strong-type diagrams are both convex, with the latter contained in the former. Furthermore, if two points lie in the former, then the open interval connecting them lies in the latter. Determining the type diagrams of various operators
is a basic task of harmonic analysis, as it conveys a lot of information as to how
transforms the width and height of functions.
There is a different interpolation method, the complex interpolation method, which offers similar results to the real interpolation method but with some slight differences. On the plus side, the complex interpolation method gives sharper bounds, and more importantly can handle the case where the operator
itself varies (analytically) with the interpolation parameter. On the minus side, the method cannot upgrade weak or restricted weak-type estimates to strong-type estimates.
In the next set of notes we shall present several applications of the real interpolation method.
— 5. Miscellaneous exercises —
Exercise 53 (Loomis-Whitney inequality) Let
, let
be measure spaces, and for
let
for some
. Show that the function 
lies in
with the Loomis-Whitney inequality 
Conclude in particular the box inequality 
where
is any subset of
and
is the canonical projection from
to
. From this, deduce the weak isoperimetric inequality 
for any
, where
is the Lebesgue measure of
and
is the
-dimensional Hausdorff measure of the boundary of
.
Exercise 54 (Borel-Cantelli lemma for functions) Let
be such that
. Show that
converges to zero pointwise almost everywhere.
Exercise 55 (Borel-Cantelli lemma for sets) Let
be such that
. Show that almost every
is contained in only finitely many of the
.