Showing posts with label calculus. Show all posts
Showing posts with label calculus. Show all posts

2026-08-29

Interpreting a Model's Statistical Significance with Many Parameters

This post is about the situation when developing statistical models, in which parameters are given with their means and standard errors, of ensuring that a data analyst does not mistakenly reject the null hypothesis for those parameters with small enough standard errors if those smaller standard errors were still ultimately the result of chance. Common correction methods for such assurance include but are not limited to the Bonferroni correction. This post discusses those correction methods as well as an alternative that I recently thought of. Follow the jump to see everything else, because even the introduction, which is meant to be a brief introduction to statistical models, is long enough that the jump would otherwise be too far below & would break the flow of this post for a reader.

2024-11-03

Some Dangers of Confusing "Changing One's Mind" with "Bayesian Updating"

Recent conversations with friends & colleagues about probability theory reminded me of conversations with a friend of mine in graduate school about the supposed virtues of making one's own reasoning in one's daily life more systematic through Bayesian inference. The basic idea, in rough qualitative terms, is that one's belief in a hypothesis can be quantified through a prior probability, and when one observes some data related to that hypothesis, one can use the probabilities of observing that data when that hypothesis does or does not hold to update one's belief in (becoming the posterior probability of) that hypothesis based on the data. An example of quantitative & qualitative explanations can be found on the site LessWrong [LINK]. However, even in graduate school and again more recently, I realized that it is very easy for one to talk oneself into believing that one is using systematic Bayesian reasoning while actually just rationalizing one's own prior beliefs & changes in beliefs after the fact. This can be illustrated mathematically in a few ways that are not exhaustive. Follow the jump to see more.

2024-05-02

Finite Determinants of Linear Operators in Continuous Vector Spaces

Recently, I wondered whether it is possible for a linear operator in a continuous (infinite-dimensional) vector space to have a finite determinant. By "continuous vector space", I mean that the identity operator can be resolved for a complete orthonormal basis \( |\phi(x) \rangle \) for all \( x \) such that \( \langle \phi(x), \phi(x') \rangle = \delta(x - x') \) as \( \hat{1} = \int |\phi(x)\rangle\langle \phi(x)|~\mathrm{d}x \). If an operator \( \hat{A} \) has continuous matrix elements \( A(x, x') = \langle \phi(x), \hat{A}\phi(x') \rangle \), then it is easy to see that the conditions for its trace \( \operatorname{trace}(\hat{A}) = \int A(x, x)~\mathrm{d}x \) to be finite are that the integral must converge, so the "function" \( A(x, x) \) must asymptotically approach 0 strictly faster than \( 1/x \) as \( |x| \to \infty \) and must at most have singularities at finite points \( x_{0} \) that diverge strictly slower than \( 1/|x - x_{0}| \). This can be seen as the continuum limit of a sum over the diagonal. However, the determinant is harder to express in this way because it involves products over diagonals & subdiagonals that are harder to express in a continuum space.

For this post, I will only consider Hermitian positive-definite operators. The conditions that I will list for which the determinant exists for such operators are sufficient for the determinant to exist, but I am not convinced that they are necessary. If such operators have an eigenvalue decomposition \( \hat{A} = \int a(x) |\phi(x)\rangle\langle \phi(x)|~\mathrm{d}x \) where the vectors \( \{ |\phi(x) \rangle \} \) form a complete orthonormal basis and the eigenvalues satisfy \( a(x) > 0 \) for all \( x \), then one can make use of the identity \( \ln(\det(\hat{A})) = \operatorname{trace}(\ln(\hat{A})) \) to say that \( \ln(\det(\hat{A})) = \int \ln(a(x))~\mathrm{d}x \). For the right-hand side to converge, then \( \ln(a(x)) \) must asymptotically approach 0 with \( x \) as \( |x| \to \infty \) strictly faster than \( 1/x \), which means that \( a(x) \) must asymptotically 1 with \( x \) as \( |x| \to \infty \) strictly faster than \( \exp(1/x) \) (which is not the same as \( e^{-x} \)), and \( \ln(a(x)) \) can at most have singularities at finite points \( x_{0} \) that diverge strictly slower than \( 1/|x - x_{0}| \), which means that \( a(x) \) must either diverge to \( \infty \) strictly slower than \( \exp(1/|x - x_{0}|) \) or drop to 0 strictly slower than \( \exp(-1/|x - x_{0}|) \). For example, \( a(x) = \exp(1/(x^{2} + x_{0}^{2})) \) fits the bill; note that this is not the same as the Gaussian kernel \( \exp(-(x^{2} + x_{0}^{2})) \). Intuitively, this condition makes sense, because for a finite-dimensional diagonal matrix as the dimension becomes arbitrarily large, the diagonal elements must mostly be exactly or very close to 1 for the determinant to not grow arbitrarily large with the dimension.

In finite-dimensional vector spaces, it is also easy to compute the determinants of triangular matrices simply as the products of the diagonal elements. (This is why the determinant is most often computed by an algorithm like first computing the LU decomposition and then taking the product of the diagonal elements of the upper-triangular matrix, which for an \( N \times N \) matrix involves \( O(N^{3}) \) operations, as opposed to the Leibniz formula involving every permutation which involves \( O(N!N) \) operations.) In infinite-dimensional vector spaces, a matrix that is triangular in a countable basis can have the determinant computed similarly as in finite-dimensional vector spaces; if an operator \( \hat{A} \) in that basis has elements \( A_{ij} \), then using the definition \( \ln(|\det(\hat{A})|) = \prod_{i} \ln(|A_{ii}|) \), the determinant converges as long as the diagonal elements \( |A_{ii}| \) are mostly exactly or very close to 1, specifically such that as \( |i| \to \infty \), \( \ln(|A_{ii}|) \) decays to 0 strictly faster than \( 1/i \). (Note that \( i \) is an integer index written in slanted font, not the imaginary unit \( \operatorname{i} \) written in upright font.) However, I am not sure how to generalize this to operators that are expressed as triangular matrices in continuous bases.

2024-04-01

Transitioning from microscopic to macroscopic and quantum to classical regimes

I recently read two things that were of interest to me having previously worked in physics. One was an article in The New Yorker magazine [LINK], in which the author does a good job of going over the successes of mathematical modeling in the physical sciences and contrasting this with the limitations of mathematical modeling in public health (showing, for example, how many models of the spread of contagions fail when governments & societies take fast & drastic collective actions to limit the spread), the failures of mathematical models in social sciences where the outputs of those models can create feedback loops with public sentiment (for example in political polling), and the way that many people who use machine learning models in different domains expect the fancy curve-fitting of those models to represent fundamental understanding when that might not really be so. The other was a journal article published in Physical Review Letters [LINK] about how it can be possible to test the extent to which a massive (as opposed to massless) object which exhibits the dynamics of a simple harmonic oscillator and prepared in a quantum coherent state can be tested for deviations from classical behavior using a protocol that does not depend on the mass of the object (although I question this given that the protocol depends on timed measurements that depend on the frequency of oscillation, and in many physics contexts the frequency does depend on the mass as \( \omega = \sqrt{k/m}\), but this is somewhat of a quibble). These two things got me to think about something that I realized I never got out of many years of formal undergraduate & graduate education in physics. This can be illustrated with the following example.

In introductory physics classes that focus on Newtonian mechanics, a prototypical problem involves a block, modeled as a point mass, sliding (with or without friction) down a fixed triangular incline in the constant gravitational field of the Earth. In the context of those classes, instructors will be careful to note that this is merely a model, and corrections could come from the inclusion of the variation of the Earth's gravitational field & surface curvature, the technical possibility of moving the triangular incline (which must be much more massive than the block in question), the shape of the block, variations in the touching surfaces, air resistance, et cetera. In later classes, instructors may point out corrections due to special relativity (i.e. the speed of light) and general relativity (as it relates to the Earth's gravitational field).

However, in later classes about quantum mechanics & statistical mechanics, instructors explain how different the models are from models of Newtonian mechanics at human scales, but they often promise that appropriate treatments of aggregates of microscopic constituents can consistently recover results from Newtonian mechanics, yet this promise is almost never fulfilled. In particular, wavefunctions that describe pure states of single microscopic particles are quite far removed from the simple dynamical variables describing blocks on inclined planes, although statistical mechanics can probabilistically describe the solid states of the block & inclined plane as well as the gaseous state of the surrounding air, it is not usually extended to describe the dynamics of the block sliding down the inclined plane. For example, if a block sliding down a fixed inclined plane of horizontal angle \( \theta \) in a uniform gravitational field is described as having equations of motion \( m\ddot{x} = mg\sin(\theta) \) where the \( x \)-axis is defined as pointing downward parallel to the slope of the inclined plane for increasing \( x \) and the \( y \)-axis points outward in the normal direction from the inclined plane, then I wish to see corrections of the form \( m\ddot{\vec{x}} = \sum_{\mu = 0}^{\infty} \sum_{\nu = 0}^{\infty} \hbar^{\mu} k_{\mathrm{B}}^{\nu} \vec{f}^{(\mu, \nu)} \) where the lowest-order term is \( \vec{f}^{(0, 0)} = mg\sin(\theta)\vec{e}_{x} \). I have never seen these sorts of quantum or statistical corrections to Newtonian equations of motion in simple (in the context of Newtonian mechanics) systems. Similarly, it is rare to see how quantum or statistical mechanical systems can, in appropriate limits, reproduce classical systems; I can only think of the quantum coherent state of the simple harmonic oscillator as well as how the Moyal bracket in the phase space formulation of quantum mechanics reduces to lowest order in \( \hbar \) to the Poisson bracket, and in the latter case, intuitive construction of the quantum phase space quasiprobability function is made more difficult (compared to construction of a classical phase space probability density function, as I did in a post [LINK] from a few years ago) by the fact that unlike the classical phase space probability density function, the quantum phase space quasiprobability function cannot be arbitrarily localized in phase space, it can take on negative values for certain wavefunctions, it is compressible in phase space with respect to its own evolution over time, and it is not obvious how it should look for a system of many particles constituting a macroscopic object like a block (in contrast to a classical phase space probability density function, which for such a system could just be a product of Dirac delta functions localizing each microscopic constituent to a point in phase space).

These considerations reminded me of a discussion I had last year with friends from college, who also did course 8 (physics) with me. We came to a consensus that while people who do not become physics majors should, as usual, get exposure to Newtonian physics and the basics of electricity & magnetism, people who become physics majors should have a curriculum over 3-4 years that exhibits a sensible conceptual progression. In particular, after seeing Newtonian mechanics, such students should then be exposed to Lagrangian & Hamiltonian formulations of classical mechanics. The Lagrangian formulation of classical mechanics should then be used to develop intuitions about mechanical waves, which in turn can lead to introductions to classical field theory and development of classical electromagnetic theory as a rich example of a classical field theory. (I would also personally recommend using the introduction of mechanical waves to introduce the linear algebraic treatment of waves and then reintroduce the linear algebraic treatment of waves into the treatment of linear classical field theories in general & linear classical electromagnetic theory in particular.) The Hamiltonian formulation of classical mechanics should then be used to develop intuitions about probability distributions in classical mechanics, which in turn can be used to develop intuitions about statistical mechanics. Optionally, at this point, the Hamiltonian formulation of classical mechanics can also be used to develop intuitions about nonlinear dynamics & chaos theory, but while this is good for the broader education of physics students, it is less immediately relevant for the introduction of quantum theory to come soon after (because quantum mechanics is linear). Finally, only after these things happen should quantum theory be introduced, such that there are clear connections of the wavefunction formulation of quantum mechanics to mechanical waves, the phase space formulation of quantum mechanics to classical phase space probability distributions, and the linear algebraic framework of quantum mechanics to linear algebraic treatments of classical field theories (including linear classical electromagnetic theory); this will ensure that students understand how ideas like superposition, interference, rotation through a Hilbert space, statistical uncertainty, and related ideas are not unique to quantum mechanics (which is unfortunately too often a consequence of the way quantum mechanics is typically introduced in undergraduate curricula, at least in the US). We also came to a consensus that in each course, there should be clear explanations of what prototypical systems are analytically solvable, what prototypical systems are not analytically solvable, and why (in each case).

2023-11-01

Contravariant and Covariant Objects in Matrix Notation

For many years when and since I was in college, I wondered whether it might be possible to consistently represent contravariant & covariant objects using vector & matrix notation. In particular, when I learned about the idea of covariant representations of [invariant] vectors being duals to contravariant representations of [invariant] vectors, meaning that if a contravariant representation of a [invariant] vector can be seen as a column vector, then a covariant representation of a [invariant] vector can be seen as a row vector, I wondered how it would be possible to represent the fully covariant metric tensor as a metric tensor if it multiplies a contravariant representation of a [invariant] vector (i.e. a column vector) to yield a covariant representation of a [invariant] vector (i.e. a row vector), especially as traditionally in linear algebra, a matrix acting on a column vector yields another column vector (while transposition, though linear in the sense of respecting addition and scalar multiplication, cannot be represented simply as the action of another matrix). At various points, I've wondered if this means that fully contravariant or fully covariant representations of multi-index tensors should be represented as columns of columns or rows of rows, and I've tried to play around with these ideas more. This post is not the first to explore such ideas even online, as I came across notes online by Viktor T. Toth [LINK], but this post is my attempt to flesh out these ideas further. Follow the jump to see more. Throughout this post, I will work with the notation of 2 spatial indices, in which the fully covariant representation of the metric tensor \( g_{ij} = \vec{e}_{i} \cdot \vec{e}_{j} \) might not be Euclidean, where indices will use English letters \( i, j, k, \ldots \in \{1, 2\} \), where superscripts do not imply exponents, and where multiple superscripts do not imply single numbers (for example, \( g_{12} \) is the fully covariant component of the metric tensor with first index 1 and second index 2, not the covariant component at index 12 of a single-index tensor (vector)); extensions to spacetime (where the convention is to use indices labeled by Greek letters) and in particular to 3 spatial + 1 temporal dimensions are trivial. Additionally, Einstein summation will be assumed, and all tensors (including vectors & scalars) are assumed to be real-valued. Finally, I will do my best to ensure that when indices are raised or lowered, the ordering of indices is clear (as examples, distinguishing \( T^{i}_{\, j} \) from \( T_{i}^{\, j} \) instead of ambiguously using \( T^{i}_{j} \) or \( T^{j}_{i} \)), but this will depend on the quality of LaTeX rendering in this post.

2022-12-02

Fundamental Theorem of Calculus for Functionals

I happened to think more about the idea of recovering a functional by somehow integrating its functional derivative. In the process, I realized that certain ideas that I would have to consider make this post a natural follow-up to a recent post [LINK] about mapping scalars to functions. This will become clear later in this post.

For a single variable, a function \( f(x) \) has an antiderivative \( F(x) \) such that \( f(x) = \frac{\mathrm{d}F}{\mathrm{d}x} \). One statement of the fundamental theorem of calculus is that this implies that \[ \int_{a}^{b} f(x)~\mathrm{d}x = F(b) - F(a) \] for these functions. In turn, this means \( F(x) \) can be extracted directly from \( f(x) \) through \[ F(x) = \int_{x_{0}}^{x} f(x')~\mathrm{d}x' \] in which \( x_{0} \) is chosen such that \( F(x_{0}) = 0 \).

For multiple variables, a conservative vector field \( \mathbf{f}(\mathbf{x}) \) in which \( \mathbf{f} \) must have the same number of components as \( \mathbf{x} \) can be said to have a scalar antiderivative \( F(\mathbf{x}) \) in the sense that \( \mathbf{f} \) is the gradient of \( F \), meaning \( \mathbf{f}(\mathbf{x}) = \nabla F(\mathbf{x}) \); more precisely, \( f_{i}(x_{1}, x_{2}, \ldots, x_{N}) = \frac{\partial F}{\partial x_{i}} \) for all \( i \in \{1, 2, \ldots, N \} \). (Note that if \( \mathbf{f} \) is not conservative, then it by definition cannot be written as the gradient of a scalar function! This is an important point to which I will return later in this post.) In such a case, a line integral (which, as I will emphasize again later in this post, is distinct from a functional path integral) from vector point \( \mathbf{a} \) to vector point \( \mathbf{b} \) of \( \mathbf{f} \) can be computed as \( \int \mathbf{f}(\mathbf{x}) \cdot \mathrm{d}\mathbf{x} = F(\mathbf{b}) - F(\mathbf{a}) \); more precisely, this equality holds along any contour, so if a contour is defined as \( \mathbf{x}(s) \) for \( s \in [0, 1] \), no matter what \( \mathbf{x}(s) \) actually is, as long as \( \mathbf{x}(0) = \mathbf{a} \) and \( \mathbf{x}(1) = \mathbf{b} \) hold, then \[ \sum_{i = 1}^{N} \int_{0}^{1} f_{i}(x_{1}(s), x_{2}(s), \ldots, x_{N}(s)) \frac{\mathrm{d}x_{i}}{\mathrm{d}s} \mathrm{d}s = F(\mathbf{b}) - F(\mathbf{a}) \] must also hold. This therefore suggests that \( F(\mathbf{x}) \) can be extracted from \( \mathbf{f}(\mathbf{x}) \) by relabeling \( \mathbf{x}(s) \to \mathbf{x}'(s) \), \( \mathbf{a} \) to a point such that \( F(\mathbf{a}) = 0 \), and \( \mathbf{b} \to \mathbf{x} \). Once again, if \( \mathbf{f}(\mathbf{x}) \) is not conservative, then it cannot be written as the gradient of a scalar field \( F \), and the integral \( \sum_{i = 1}^{N} \int_{0}^{1} f_{i}(x_{1}(s), x_{2}(s), \ldots, x_{N}(s)) \frac{\mathrm{d}x_{i}}{\mathrm{d}s} \mathrm{d}s \) will depend on the specific choice of \( \mathbf{x}(s) \), not just the endpoints \( \mathbf{a} \) and \( \mathbf{b} \).

For continuous functions, the generalization of a vector \( \mathbf{x} \), or more precisely \( x_{i} \) for \( i \in \{1, 2, \ldots, N\} \), is a function \( x(t) \) where \( t \) is a continuous dummy index or parameter analogous to the discrete index \( i \). This means the generalization of a scalar field \( F(\mathbf{x}) \) is the scalar functional \( F[x] \). What is the generalization of a vector field \( \mathbf{f}(\mathbf{x}) \)? To be precise, a vector field is a collection of functions \( f_{i}(x_{1}, x_{2}, \ldots, x_{N}) \) for all \( i \in \{1, 2, \ldots, N \} \). This suggests that its generalization should be a function of \( t \) and must somehow depend on \( x(t) \) as well. It is tempting therefore to write this as \( f(t, x(t)) \) for all \( t \). However, although this is a valid subset of the generalization, it is not the whole generalization, because vector fields of the form \( f_{i}(x_{i}) \) are collections of single-variable functions that do not fully capture all vector fields of the form \( f_{i}(x_{1}, x_{2}, \ldots, x_{N}) \) for all \( i \in \{1, 2, \ldots, N \} \). As a specific example, for \( N = 2 \), the vector field with components \( f_{1}(x_{1}, x_{2}) = (x_{1} - x_{2})^{2} \) and \( f_{2}(x_{1}, x_{2}) = (x_{1} + x_{2})^{3} \) cannot be written as just \( f_{1}(x_{1}) \) and \( f_{2}(x_{2}) \), as \( f_{1} \) depends on \( x_{2} \) and \( f_{2} \) depends on \( x_{1} \) as well. Similarly, in the generalization, one could imagine a function of the form \( f = \frac{x(t)}{x(t - t_{0})} \mathrm{exp}(-(t - t_{0})^{2}) \); in this case, it is not correct to write it as \( f(t, x(t)) \) because the dependence of \( f \) on \( x \) at a given dummy index value \( t \) comes through not only \( x(t) \) but also \( x(t - t_{0}) \) for some fixed parameter \( t_{0} \). Additionally, the function may depend not only on \( x \) per se but also on derivatives \( \frac{\mathrm{d}^{n} x}{\mathrm{d}t^{n}} \); the case of the first derivative \( \frac{\mathrm{d}x}{\mathrm{d}t} = \lim_{t_{0} \to 0} \frac{x(t) - x(t - t_{0})}{t_{0}} \) illustrates the connection to the aforementioned example. Therefore, the most generic way to write such a function is effectively as a functional \( f[x; t] \) with a dummy index \( t \). The example \( f = \frac{x(t)}{x(t - t_{0})} \mathrm{exp}(-(t - t_{0})^{2}) \) can be formalized as \( f[t, x] = \int_{-\infty}^{\infty} \frac{x(t')}{x(t' - t_{0})} \mathrm{exp}(-(t' - t_{0})^{2}) \delta(t - t')~\mathrm{d}t' \) where the dummy index \( t' \) is the integration variable while the dummy index \( t \) is free. (For \( N = 3 \), the condition of a vector field being conservative is often written as \( \nabla \times \mathbf{f}(\mathbf{x}) = 0 \). I have not used that condition in this post because the curl operator does not easily generalize to \( N \neq 3 \).)

If a functional \( f[x; t] \) is conservative, then there exists a functional \( F[x] \) (with no free dummy index) such that \( f \) is the functional derivative \( f[x; t] = \frac{\delta F}{\delta x(t)} \). Comparing the notation between scalar fields and functionals, \( \sum_{i} A_{i} \to \int A(t)~\mathrm{d}t \) and \( \mathrm{d}x_{i} \to \delta x(t) \), in which \( \delta x(t) \) is a small variation in a function \( x \) specifically at the index value \( t \) and nowhere else. This suggests a generalization of the fundamental theorem of calculus to functionals as follows. If \( a(t) \) and \( b(t) \) are fixed functions, then \( \int_{-\infty}^{\infty} \int f[x; t]~\delta x(t)~\mathrm{d}t = F[b] - F[a] \). More precisely, a path from the function \( a(t) \) to the function \( b(t) \) at every index value \( t \) can be parameterized by \( s \in [0, 1] \) by the map \( s \to x(t, s) \) which is a function of \( t \) for each \( s \) such that \( x(t, 0) = a(t) \) and \( x(t, 1) = b(t) \); this is why I linked this post to the most recent post on this blog. With this in mind, the fundamental theorem of calculus becomes \[ \int_{-\infty}^{\infty} \int_{0}^{1} f[x(s); t] \frac{\partial x}{\partial s}~\mathrm{d}s~\mathrm{d}t = F[b] - F[a] \] where, in the integrand, the argument \( x \) in \( f \) has the parameter \( s \) explicit but the dummy index \( t \) implicit; the point is that this equality holds regardless of the specific parameterization \( x(t, s) \) as long as \( x \) at the endpoints of \( s \) satisfies \( x(t, 0) = a(t) \) and \( x(t, 1) = b(t) \). This also means that \( F[x] \) can be recovered if \( b(t) = x(t) \) and \( a(t) \) is chosen such that \( F[a] = 0 \), in which case \[ F[x] = \int_{-\infty}^{\infty} \int_{0}^{1} f[x'(s); t]~\frac{\partial x'}{\partial s}~\mathrm{d}s~\mathrm{d}t \] (where \( x(t, s) \) has been renamed to \( x'(t, s) \) to avoid confusion with \( x(t) \)). If \( f[x; t] \) is not conservative, then there is no functional \( F[x] \) whose functional derivative with respect to \( x(t) \) would yield \( f[x; t] \); in that case, with \( x(t, 0) = a(t) \) and \( x(t, 1) = b(t) \), the integral \( \int_{-\infty}^{\infty} \int_{0}^{1} f[x(s); t] \frac{\partial x}{\partial s}~\mathrm{d}s~\mathrm{d}t \) does depend on the specific choice of parameterization \( x(t, s) \) with respect to \( s \) and not just on the functions \( a(t) \) and \( b(t) \) at the endpoints of \( s \).

As an example, consider from a previous post [LINK] the nonrelativistic Newtonian action \[ S[x] = \int_{-\infty}^{\infty} \left(\frac{m}{2} \left(\frac{\mathrm{d}x}{\mathrm{d}t}\right)^{2} + F_{0} x(t) \right)~\mathrm{d}t \] for a particle under the influence of a uniform force \( F_{0} \) (which may vanish). The first functional derivative is \[ f[x; t] = \frac{\delta S}{\delta x(t)} = F_{0} - m\frac{\mathrm{d}^{2} x}{\mathrm{d}t^{2}} \] and its vanishing would yield the usual equation of motion. The action itself vanishes for \( x(t) = 0 \), which will be helpful when using the fundamental theorem of calculus to recover the action from the equation of motion. In particular, one can parameterize \( x'(t, s) = sx(t) \) such that \( x'(t, 0) = 0 \) and \( x'(t, 1) = x(t) \). This gives the integral \( \int_{0}^{1} \left(F_{0} - ms\frac{\mathrm{d}^{2} x}{\mathrm{d}t^{2}}\right)x(t)~\mathrm{d}s = F_{0} x(t) - \frac{m}{2} x(t) \frac{\mathrm{d}^{2} x}{\mathrm{d}t^{2}} \). This is then integrated over all \( t \), so the first term is identical to the corresponding term in the definition of \( S[x] \), and the second term becomes the same as the corresponding term in the definition of \( S[x] \) after integrating over \( t \) by parts and setting the boundary conditions that \( x(t) \to 0 \) for \( |t| \to \infty \). (Other boundary conditions may require more care.) In any case, the parameterization \( x'(t, s) = sx(t) \) is not the only choice that could fulfill the boundary conditions; the salient point is that any parameterization fulfilling the boundary conditions would yield the correct action \( S[x] \).

I considered that example because I wondered whether any special formulas need to be considered if \( f[x; t] \) depends explicitly on first or second derivatives of \( x(t) \), as might be the case in nonrelativistic Newtonian mechanics. That example shows that no special formulas are needed because even if the Lagrangian explicitly depends on the velocity \( \frac{\mathrm{d}x}{\mathrm{d}t} \), the action \( S \) only explicitly depends as a functional on \( x(t) \), so proper application of functional differentiation and regular integration by parts will ensure proper accounting of each piece.

This post has been about the fundamental theorem of calculus saying that the 1-dimensional integral of a function in \( N \) dimensions along a contour, if that function is conservative, is equal to the difference between the two endpoints of its scalar antiderivative. This generalizes easily to infinite dimensions and continuous functions instead of finite-dimensional vectors. There is another fundamental theorem of calculus saying that the \( N \)-dimensional integral in a finite volume of the scalar divergence of an \( N \)-dimensional vector function, if that volume has a closed orientable surface, is equal to the \( N - 1 \)-dimensional integral of the inner product of that function with the normal vector (of unit 2-norm) at every point on the surface across the whole surface, meaning \[ \int_{V} \sum_{i = 1}^{N} \frac{\partial f_{i}}{\partial x_{i}}~\mathrm{d}V = \oint_{\partial V} \sum_{i = 1}^{N} f_{i}(x_{1}, x_{2}, \ldots, x_{N}) n_{i}(x_{1}, x_{2}, \ldots, x_{N})~\mathrm{d}S \] where \( \sum_{i = 1}^{N} |n_{i}(x_{1}, x_{2}, \ldots, x_{N})|^{2} = 1 \) for every \( \mathbf{x} \). From a purely formal perspective, this could generalize to something like \( \int_{V} \int_{-\infty}^{\infty} \frac{\delta f[x; t]}{\delta x(t)}~\mathrm{d}t~\mathcal{D}x = \oint_{\partial V} \int_{-\infty}^{\infty} f[x; t]n[x; t]~\mathrm{d}t~\mathcal{D}x \) having generalized \( \frac{\partial}{\partial x_{i}} \to \frac{\delta}{\delta x(t)} \), \( \prod_{i} \mathrm{d}x_{i} \to \mathcal{D}x \), and \( n_{i}(\mathbf{x}) \to n[x; t] \) where \( n[x; t] \) is normalized such that \( \int_{-\infty}^{\infty} |n[x; t]|^{2}~\mathrm{d}t = 1 \) for all \( x(t) \) on the surface. However, this formalism may be hard to further develop because the space has infinite dimensions. Even when working in a countable basis, it might not be possible to characterize an orientable surface enclosing a volume in an infinite-dimensional space; the surface is also infinite-dimensional. While the choice of basis is arbitrary, things become even less intuitive when choosing to work in an uncountable basis.

2022-11-01

Mapping Scalars to Functions

In just over a year, I've written three posts for this blog about functionals, specifically about their application to probability theory [LINK], finding their stationary points [LINK], and the use of their stationary points in classical mechanics [LINK]. As a reminder, a functional is an object that maps a space of functions to a space of numbers. This got me thinking about what the reverse, namely an object that maps a space of numbers to a space of functions, looks like. To be clear, this is not the same as an ordinary function which, as an element in a space of functions, maps a space of numbers to a space of numbers.

As I thought about it more, I realized that this is a bit easier to understand and therefore more commonly encountered than a functional. An extremely glib way to describe such an object is a function of multiple variables. However, it may be more enlightening to describe this in further detail to avoid potentially deceptive images that may arise from that glib description.

In the discrete case, the matrix elements \( A_{ij} \) can be described as a map from integers to vectors, in which an integer \( j \) is associated with a vector whose elements indexed by an integer \( i \) are \( A_{ij} \). This is the essential idea behind seeing the columns of the matrix with elements \( A_{ij} \) as a collection of vectors. Formally, this maps \( i \to (j \to A_{ij}) \) where the map \( j \to A_{ij} \) defines a vector indexed by the free variable \( i \).

Similarly, in the continuous case, the function elements \( f(x, y) \) can be described as a map from numbers to functions, in which a number \( y \) is associated with a function whose elements indexed by a number \( x \) are \( f(x, y) \). Formally, this maps \( x \to (y \to f(x, y)) \) where the map \( y \to f(x, y) \) defines a function indexed by the free variable \( x \). These ideas are foundational to the development of more abstract notions of functions, like lambda calculus.

2017-12-18

Book Review: "Hidden Figures" by Margot Lee Shetterly

I've recently read the book Hidden Figures by Margot Lee Shetterly. It weaves together the true stories of a few particular mathematicians, who happened to be black women (among a larger group of such female black mathematicians), who made extremely important contributions to the development of American warplanes in WWII and then spacecrafts during the 1950s and 1960s, including the crafts that took John Glenn to space and then the Apollo 11 astronauts to the moon. It highlights the skills of these women and their own personal lives, in conjunction with the broader social issues of that time.

The book is moderately long, but it is very well-written and engaging. I liked seeing the descriptions of these towns that flourished during the wartime years and the space race as bustling with life and energy, because with the trends of deindustrialization starting from a few decades ago, I haven't really been able to see descriptions such towns as much beyond shells of their former selves. This also ties in with the discussions of the military-industrial complex and how the formation of these towns during the wartime was a symptom of that phenomenon, which in turn meant that as wartime research facilities and organizations were often temporary, even if they hired black people, allowing them to economically advance to the middle class, those economic advancements became tenuous due to the temporary nature of such jobs, such that when the goal was met (whether it was winning the war or landing on the moon), those facilities would be closed and the employees there would be displaced with few, if any, alternatives available to them. Of course, pervading the book were descriptions of the explicit and implicit forms of institutionalized racism and sexism, whether at work in the form of barriers to career advancement or collegiality/free exchange of ideas, or in the context of daily life with respect to the civil rights movement, sit-ins, et cetera. Not only were those issues discussed in a broader context, but their impact on the specific protagonists of the book was detailed, showing how these women had to deal with so many struggles just to stay afloat while still trying to achieve the same goals to which any other family of any ethnicity would strive, namely, caring for spouses and children, putting food on the table, balancing work and family, and being able to raise children in a safe environment and educate them well; it really helped that the author so masterfully portrayed the mundanity of daily daily life for these women to show how stupid obstacles, like legalized segregation and institutionalized barriers to career advancement, could get in the way of the passion that these women had for STEM. It was also interesting to see that black communities like those in this book were acutely aware of how much more advanced the USSR and other communist countries were in terms of race and gender relations, and actively called out the US on its own failings in that regard (in the context of the US trying to ally with African and Asian countries that used to be European colonies), while these black women, despite being in the middle of such institutional bigotry, kept their heads held high and persevered in pursuit of their goals to contribute to STEM R&D. Related to that, it was also chilling to see how de facto segregation has persisted in education in many places through the US resulting in school facilities that in many poor places are no better than they were several decades ago, and also to see how many of the arguments for white parents sending their kids to private schools at that time were more explicitly about preserving racial segregation in education. Overall, I enjoyed this book thoroughly and would strongly recommend it to people for a clear and engaging account of how NASA and the social issues of the middle of the 20th century became intertwined.

2010-10-04

Isaac Newton, Progress, and Patents

In my physics recitation class today, our recitation leader briefly digressed from the material at hand to discuss the history of differential calculus and the conflict between Isaac Newton and Gottfried Leibniz. Basically, Newton claimed to have invented differential calculus first (although, as with any other "invention", neither can truly claim to have invented calculus from scratch as they were building on the work of mathematicians before them (and I don't just mean 1 + 1 = 2 — I mean things like infinite series and tangent lines)), but as he kept his work secret for decades, he ended up publishing his work on calculus after Leibniz published his work. While both were initially on good terms, as Newton became more possessive of his own work and convinced of his own originality, the debate became progressively more heated, with Newton and his supporters accusing Leibniz of plagiarism. Follow the jump to read more.