Showing posts with label mathematics. Show all posts
Showing posts with label mathematics. Show all posts

2026-08-29

Interpreting a Model's Statistical Significance with Many Parameters

This post is about the situation when developing statistical models, in which parameters are given with their means and standard errors, of ensuring that a data analyst does not mistakenly reject the null hypothesis for those parameters with small enough standard errors if those smaller standard errors were still ultimately the result of chance. Common correction methods for such assurance include but are not limited to the Bonferroni correction. This post discusses those correction methods as well as an alternative that I recently thought of. Follow the jump to see everything else, because even the introduction, which is meant to be a brief introduction to statistical models, is long enough that the jump would otherwise be too far below & would break the flow of this post for a reader.

2025-10-02

FOLLOW-UP: Learning and Making Sense of Differential Geometry in General Relativity

This post is a follow-up to the previous post [LINK] in which I tried to make sense of the mathematics of differential geometry, especially in the context of general relativity, and proposed notation that may be less confusing, more consistent, and arguably more powerful than traditional notation given the existence of a metric in general relativity. With similar motivations, this post explores how some of the ideas of electromagnetic (EM) theory may change on curved manifolds. Follow the jump to see more.

2025-09-14

Learning and Making Sense of Differential Geometry in General Relativity

I have been out of the field of physics for over 5 years, didn't get to make use of my training in quantitative analysis too much in my previous job as a postdoctoral researcher at UC Davis, and make even less use of my training in quantitative analysis in my current job as a transportation planner at Cambridge Systematics. When I was doing work in my PhD in physics and when I was a student in school and college before graduate school, I enjoyed solving problems in math & physics and growing & applying my toolbox of quantitative skills, so since leaving the field, and especially more recently, I have felt a slight itch to recover some of those skills purely for my own personal satisfaction. To that end, I've resolved to learn or relearn some parts of math that I did not learn at all or to the full extent that I should have in college or graduate school (because I could ultimately manage my work in my PhD without knowing those things to that extent). Currently, I'm interested in learning differential geometry as well as complex analysis. It remains to be seen whether there are other topics in physics-relevant math that I become interested in; if they are, then I will certainly make an effort to learn at least a little bit about them. I've enjoyed learning these topics in math to broaden my knowledge & skills, and I've particularly enjoyed pondering definitions & rules of these concepts in math almost like a lawyer (which also makes this useful for my current job in a very indirect way, because my current job in part involves analyzing laws & regulations and creating intuitive explanations of them for public sector clients).

I can't guarantee that I'll write a blog post about every topic in math that I learn about. However, I am writing this post specifically about differential geometry in parts because I feel that I have learned the basic ideas in it in the context of general relativity to my satisfaction (which was my original goal) and because I have some questions/concerns that I have not been able to satisfactorily resolve based on what I have read in lecture notes or textbooks. This post is a way to further flesh out those questions/concerns. Follow the jump to see more.

There are a few conventions & assumptions to note throughout this post.

  • The dimension of the manifold will generically be denoted \( N \). For most nonrelativistic physics, \( N = 3 \), while for most relativistic physics (including general relativity), \( N = 4 \); exceptions include constrained low-dimensional systems or different physical models that have different dimensions analogous to differential geometry.
  • I will consistently use Einstein summation unless otherwise specified. This means that indices that are repeated with one as an upper index and one as a lower index will be summed; indices should never be present more than twice at all and more than once in the same (upper versus lower) position.
  • Although the convention in general relativity is to use lowercase Greek letters for indices, I will keep things easier to read by using lowercase English letters for indices.
  • The motivation of general relativity means that I will only consider differentiable manifolds and specifically smooth (infinitely differentiable) scalar functions or tensor components on them.
  • The motivation of general relativity also means that I will only consider torsion-free connection coefficients that can be expressed in terms of the metric.

2025-01-02

Book Review: "Thinking, Fast and Slow" by Daniel Kahneman

I started reading the book Thinking, Fast and Slow by Daniel Kahneman in early 2024. This was initially recommended to me by a friend, and I became even more motivated to read it upon hearing positive things about it from colleagues at my previous job, as many of the subtleties described in the book are extremely relevant to the appropriate design of interviews, focus groups, and surveys of human subjects in social science research. However, because it is a long book and the middle of 2024 was made busier for me by moving back to Maryland, traveling a lot, and starting a new job (some of which I have discussed in a previous post [LINK]), I could not finish reading this book until much more recently. Because of this large gap between reading the initial 60% and remaining 40% of this book, I admit that I have since forgotten many details from the initial 60% of this book. Moreover, I started making notes to myself in this post based on that initial 60% because I assumed that I would be able to finish reading the remaining 40% soon afterwards and I would therefore remember the book as a coherent whole, but because that didn't happen, many of the notes that I have made in this post that were supposed to form the skeleton of this post now no longer make as much sense to me. For these reasons, this post may seem a bit more stilted than other book review posts in this blog and will likely seem stronger/more coherent when discussing the latter 40% of the book.

The book is a lengthy exposition of novel ideas in psychology & behavioral economics that were empirically validated by the author, most often in conjunction with his longtime academic collaborator Amos Tversky. The concluding chapter does a good job of recapitulating the main ideas of the book. Most of the book explores various facets of individual & group-based human behavior based on the idea that there are effectively 2 modes through which individuals process information, which the author refers to as Systems 1 & 2. System 1 "thinks fast", making snap judgments based on limited information, heuristics, and a bit of laziness, and is the aspect of thinking that drives most day-to-day reactions & decisionmaking, while System 2 "thinks slow", making more deliberate judgments with more of an effort to gather all relevant information but must in turn be consciously engaged and ultimately disengages from mental fatigue (in favor of System 1) if engaged for too long. The book also considers how individuals' typical behaviors when faced with outcomes that are certain competing with outcomes that have known or unknown probabilities deviate from behaviors idealized by microeconomic theories of expected utility, notably that while the commonly observed behavior choosing a certain gain with a lower value than the expected value of an uncertain gain can be explained to some degree by expected utility theory, the commonly observed behavior of choosing a gamble on losing outcomes with an expected loss of larger magnitude than a different certain loss cannot be explained by expected utility theory; this partly explains the risks that people take in business and can be explained in turn by how people in their perceptions tend to overestimate probabilities that are close to but not exactly 0 and underestimate probabilities that are close to but not exactly 1. Finally, the book partly explains notions of hedonic adaptation (the idea that one's sense of well-being is generally similar in many different good or bad medium- or long-term circumstances by adapting to those circumstances) by distinguishing how people rate pleasure or pain when experiencing those things versus in hindsight and shows how people's conceptions of their identities & well-being in the past, present, and future are intimately tied to their actual memories and their abilities to form & retain memories. These aspects of self-conception as well as perceptions of probability can also be tied to Systems 1 versus 2, as many seemingly shortsighted decisions or perceptions can be explained by System 1 making snap judgments lazily & using heuristics based on incomplete information.

Especially as I read the latter 40% of the book, I came to appreciate how many of the ideas of this book had permeated into other things that I had read & heard from others and that I had internalized into my own worldview & view of myself. Professionally, I could see how so many aspects of framing could be important when designing surveys & focus groups. Personally, I could see how especially as I have aged, I have in many cases consciously chosen to not worry too much about certain details and instead make decisions based on lazier heuristics because I didn't feel that the results of spending more mental energy making a decision based on System 2 would be worth the effort. At the same time, I have become more consciously aware of how my memories of things in my own life can be affected by the passage of time and by more recent events in my own life, and I have become more consciously aware of the deep entanglement between my perceptions of my own memories and the narratives that shape my perceptions of my own life & of the world. I thus feel more proud of maintaining detailed personal diaries where I take note (using System 2 as much as possible when considering things outside of the current moment) of how I feel about various things in the moment as well as in hindsight and carefully consider how & why my thoughts & feelings about different events in or aspects of my life have evolved over time. Moreover, I have become more aware over time of when I might be vulnerable (through System 1) to the power of suggestion or to a subconscious desire to align with groupthink, though given that it is System 1, I am not necessarily aware of these things until later (thinking about these things through System 2). Finally, especially over the last several years, I have come to see many things at a very broad conceptual/philosophical level, whether the experiences in my own life, the evolution of different aspects of human society, or the expansion of human knowledge, in terms of perdurantism [LINK from Wikipedia]; although I am not philosophically sophisticated enough to be able to think through & defend all of its implications, it intuitively makes sense to me to think about personal identities, feelings, people, and other things that can be said to exist, in terms of their existence in spacetime and not just in space at specific instants of time. Because of my philosophical inclination in this way, I was particularly pleased to see the author discuss the idea of time-integrated pleasure or pain and of looking at changing identities or overall life courses in terms of spacetime.

Although this book is not technical at the level of an academic journal article, it is fairly technical compared to most nonfiction books aimed at the general public, so I would say that it is aimed at a well-educated reader. That said, I do think that it is written with reasonable clarity for non-academic audiences. Additionally, the book covers many topics, and it is recommended to bear in mind the headings of sections that comprise groups of chapters, because otherwise, it is easy to lose track of the narrative of the book, especially because the book is long enough that I suspect that it would be impossible for most readers (even those who read books, including more technical nonfiction books, relatively quickly) to finish this book in one sitting. I would say that the concluding chapter is a nice way to reinforce the main points of the book in the reader's mind and that the details of each chapter can be treated as a reference when needed as opposed to forming a perfectly coherent narrative in the progression of chapters in the book.

It is important to remember that some aspects of this book are out of date. In some cases, that is just because this book was published in 2011 and had been written over many years before that; for example, the author gives an example of estimating the likelihood of choosing a particular major in college, but that example uses base rates that seem to be quite out-of-date. In other cases, the book is out of date because it is based on academic experimental work in psychology & behavioral economics, and other studies may find contradictory (either null or opposite) results to those presented in this book. The Wikipedia article about this book [LINK] discussed how most of the results from most of the studies discussed in one chapter (as an example) have been found to be not replicable, with the author afterwards admitting to putting too much faith in those studies and therefore falling prey to the same biases as those discussed in that chapter & elsewhere in the book. As a slightly different example, later parts of the book discuss the ideas of nudge theory and its seeming successes in public policy, but the Wikipedia article about nudge theory [LINK] has pointed out that later studies & meta-analyses have found that after correcting for publication biases in favor of positive results & against null results, nudging does not yield statistically significant (non-null) effects on human behavior; in this case, one of the primary researchers (who is named in this book as a collaborator of the author & pioneer of nudge theory) has made some counterarguments that I don't find convincing.

With these caveats in mind, I would still recommend this book to anyone interested in these ideas and with the patience to carefully consider them, though this may partly reflect my own biases in how I view issues of identity & the world. Follow the jump to see my other assorted & disjointed thoughts about this book.

2024-11-03

Some Dangers of Confusing "Changing One's Mind" with "Bayesian Updating"

Recent conversations with friends & colleagues about probability theory reminded me of conversations with a friend of mine in graduate school about the supposed virtues of making one's own reasoning in one's daily life more systematic through Bayesian inference. The basic idea, in rough qualitative terms, is that one's belief in a hypothesis can be quantified through a prior probability, and when one observes some data related to that hypothesis, one can use the probabilities of observing that data when that hypothesis does or does not hold to update one's belief in (becoming the posterior probability of) that hypothesis based on the data. An example of quantitative & qualitative explanations can be found on the site LessWrong [LINK]. However, even in graduate school and again more recently, I realized that it is very easy for one to talk oneself into believing that one is using systematic Bayesian reasoning while actually just rationalizing one's own prior beliefs & changes in beliefs after the fact. This can be illustrated mathematically in a few ways that are not exhaustive. Follow the jump to see more.

2024-05-02

Finite Determinants of Linear Operators in Continuous Vector Spaces

Recently, I wondered whether it is possible for a linear operator in a continuous (infinite-dimensional) vector space to have a finite determinant. By "continuous vector space", I mean that the identity operator can be resolved for a complete orthonormal basis \( |\phi(x) \rangle \) for all \( x \) such that \( \langle \phi(x), \phi(x') \rangle = \delta(x - x') \) as \( \hat{1} = \int |\phi(x)\rangle\langle \phi(x)|~\mathrm{d}x \). If an operator \( \hat{A} \) has continuous matrix elements \( A(x, x') = \langle \phi(x), \hat{A}\phi(x') \rangle \), then it is easy to see that the conditions for its trace \( \operatorname{trace}(\hat{A}) = \int A(x, x)~\mathrm{d}x \) to be finite are that the integral must converge, so the "function" \( A(x, x) \) must asymptotically approach 0 strictly faster than \( 1/x \) as \( |x| \to \infty \) and must at most have singularities at finite points \( x_{0} \) that diverge strictly slower than \( 1/|x - x_{0}| \). This can be seen as the continuum limit of a sum over the diagonal. However, the determinant is harder to express in this way because it involves products over diagonals & subdiagonals that are harder to express in a continuum space.

For this post, I will only consider Hermitian positive-definite operators. The conditions that I will list for which the determinant exists for such operators are sufficient for the determinant to exist, but I am not convinced that they are necessary. If such operators have an eigenvalue decomposition \( \hat{A} = \int a(x) |\phi(x)\rangle\langle \phi(x)|~\mathrm{d}x \) where the vectors \( \{ |\phi(x) \rangle \} \) form a complete orthonormal basis and the eigenvalues satisfy \( a(x) > 0 \) for all \( x \), then one can make use of the identity \( \ln(\det(\hat{A})) = \operatorname{trace}(\ln(\hat{A})) \) to say that \( \ln(\det(\hat{A})) = \int \ln(a(x))~\mathrm{d}x \). For the right-hand side to converge, then \( \ln(a(x)) \) must asymptotically approach 0 with \( x \) as \( |x| \to \infty \) strictly faster than \( 1/x \), which means that \( a(x) \) must asymptotically 1 with \( x \) as \( |x| \to \infty \) strictly faster than \( \exp(1/x) \) (which is not the same as \( e^{-x} \)), and \( \ln(a(x)) \) can at most have singularities at finite points \( x_{0} \) that diverge strictly slower than \( 1/|x - x_{0}| \), which means that \( a(x) \) must either diverge to \( \infty \) strictly slower than \( \exp(1/|x - x_{0}|) \) or drop to 0 strictly slower than \( \exp(-1/|x - x_{0}|) \). For example, \( a(x) = \exp(1/(x^{2} + x_{0}^{2})) \) fits the bill; note that this is not the same as the Gaussian kernel \( \exp(-(x^{2} + x_{0}^{2})) \). Intuitively, this condition makes sense, because for a finite-dimensional diagonal matrix as the dimension becomes arbitrarily large, the diagonal elements must mostly be exactly or very close to 1 for the determinant to not grow arbitrarily large with the dimension.

In finite-dimensional vector spaces, it is also easy to compute the determinants of triangular matrices simply as the products of the diagonal elements. (This is why the determinant is most often computed by an algorithm like first computing the LU decomposition and then taking the product of the diagonal elements of the upper-triangular matrix, which for an \( N \times N \) matrix involves \( O(N^{3}) \) operations, as opposed to the Leibniz formula involving every permutation which involves \( O(N!N) \) operations.) In infinite-dimensional vector spaces, a matrix that is triangular in a countable basis can have the determinant computed similarly as in finite-dimensional vector spaces; if an operator \( \hat{A} \) in that basis has elements \( A_{ij} \), then using the definition \( \ln(|\det(\hat{A})|) = \prod_{i} \ln(|A_{ii}|) \), the determinant converges as long as the diagonal elements \( |A_{ii}| \) are mostly exactly or very close to 1, specifically such that as \( |i| \to \infty \), \( \ln(|A_{ii}|) \) decays to 0 strictly faster than \( 1/i \). (Note that \( i \) is an integer index written in slanted font, not the imaginary unit \( \operatorname{i} \) written in upright font.) However, I am not sure how to generalize this to operators that are expressed as triangular matrices in continuous bases.

2024-04-01

Transitioning from microscopic to macroscopic and quantum to classical regimes

I recently read two things that were of interest to me having previously worked in physics. One was an article in The New Yorker magazine [LINK], in which the author does a good job of going over the successes of mathematical modeling in the physical sciences and contrasting this with the limitations of mathematical modeling in public health (showing, for example, how many models of the spread of contagions fail when governments & societies take fast & drastic collective actions to limit the spread), the failures of mathematical models in social sciences where the outputs of those models can create feedback loops with public sentiment (for example in political polling), and the way that many people who use machine learning models in different domains expect the fancy curve-fitting of those models to represent fundamental understanding when that might not really be so. The other was a journal article published in Physical Review Letters [LINK] about how it can be possible to test the extent to which a massive (as opposed to massless) object which exhibits the dynamics of a simple harmonic oscillator and prepared in a quantum coherent state can be tested for deviations from classical behavior using a protocol that does not depend on the mass of the object (although I question this given that the protocol depends on timed measurements that depend on the frequency of oscillation, and in many physics contexts the frequency does depend on the mass as \( \omega = \sqrt{k/m}\), but this is somewhat of a quibble). These two things got me to think about something that I realized I never got out of many years of formal undergraduate & graduate education in physics. This can be illustrated with the following example.

In introductory physics classes that focus on Newtonian mechanics, a prototypical problem involves a block, modeled as a point mass, sliding (with or without friction) down a fixed triangular incline in the constant gravitational field of the Earth. In the context of those classes, instructors will be careful to note that this is merely a model, and corrections could come from the inclusion of the variation of the Earth's gravitational field & surface curvature, the technical possibility of moving the triangular incline (which must be much more massive than the block in question), the shape of the block, variations in the touching surfaces, air resistance, et cetera. In later classes, instructors may point out corrections due to special relativity (i.e. the speed of light) and general relativity (as it relates to the Earth's gravitational field).

However, in later classes about quantum mechanics & statistical mechanics, instructors explain how different the models are from models of Newtonian mechanics at human scales, but they often promise that appropriate treatments of aggregates of microscopic constituents can consistently recover results from Newtonian mechanics, yet this promise is almost never fulfilled. In particular, wavefunctions that describe pure states of single microscopic particles are quite far removed from the simple dynamical variables describing blocks on inclined planes, although statistical mechanics can probabilistically describe the solid states of the block & inclined plane as well as the gaseous state of the surrounding air, it is not usually extended to describe the dynamics of the block sliding down the inclined plane. For example, if a block sliding down a fixed inclined plane of horizontal angle \( \theta \) in a uniform gravitational field is described as having equations of motion \( m\ddot{x} = mg\sin(\theta) \) where the \( x \)-axis is defined as pointing downward parallel to the slope of the inclined plane for increasing \( x \) and the \( y \)-axis points outward in the normal direction from the inclined plane, then I wish to see corrections of the form \( m\ddot{\vec{x}} = \sum_{\mu = 0}^{\infty} \sum_{\nu = 0}^{\infty} \hbar^{\mu} k_{\mathrm{B}}^{\nu} \vec{f}^{(\mu, \nu)} \) where the lowest-order term is \( \vec{f}^{(0, 0)} = mg\sin(\theta)\vec{e}_{x} \). I have never seen these sorts of quantum or statistical corrections to Newtonian equations of motion in simple (in the context of Newtonian mechanics) systems. Similarly, it is rare to see how quantum or statistical mechanical systems can, in appropriate limits, reproduce classical systems; I can only think of the quantum coherent state of the simple harmonic oscillator as well as how the Moyal bracket in the phase space formulation of quantum mechanics reduces to lowest order in \( \hbar \) to the Poisson bracket, and in the latter case, intuitive construction of the quantum phase space quasiprobability function is made more difficult (compared to construction of a classical phase space probability density function, as I did in a post [LINK] from a few years ago) by the fact that unlike the classical phase space probability density function, the quantum phase space quasiprobability function cannot be arbitrarily localized in phase space, it can take on negative values for certain wavefunctions, it is compressible in phase space with respect to its own evolution over time, and it is not obvious how it should look for a system of many particles constituting a macroscopic object like a block (in contrast to a classical phase space probability density function, which for such a system could just be a product of Dirac delta functions localizing each microscopic constituent to a point in phase space).

These considerations reminded me of a discussion I had last year with friends from college, who also did course 8 (physics) with me. We came to a consensus that while people who do not become physics majors should, as usual, get exposure to Newtonian physics and the basics of electricity & magnetism, people who become physics majors should have a curriculum over 3-4 years that exhibits a sensible conceptual progression. In particular, after seeing Newtonian mechanics, such students should then be exposed to Lagrangian & Hamiltonian formulations of classical mechanics. The Lagrangian formulation of classical mechanics should then be used to develop intuitions about mechanical waves, which in turn can lead to introductions to classical field theory and development of classical electromagnetic theory as a rich example of a classical field theory. (I would also personally recommend using the introduction of mechanical waves to introduce the linear algebraic treatment of waves and then reintroduce the linear algebraic treatment of waves into the treatment of linear classical field theories in general & linear classical electromagnetic theory in particular.) The Hamiltonian formulation of classical mechanics should then be used to develop intuitions about probability distributions in classical mechanics, which in turn can be used to develop intuitions about statistical mechanics. Optionally, at this point, the Hamiltonian formulation of classical mechanics can also be used to develop intuitions about nonlinear dynamics & chaos theory, but while this is good for the broader education of physics students, it is less immediately relevant for the introduction of quantum theory to come soon after (because quantum mechanics is linear). Finally, only after these things happen should quantum theory be introduced, such that there are clear connections of the wavefunction formulation of quantum mechanics to mechanical waves, the phase space formulation of quantum mechanics to classical phase space probability distributions, and the linear algebraic framework of quantum mechanics to linear algebraic treatments of classical field theories (including linear classical electromagnetic theory); this will ensure that students understand how ideas like superposition, interference, rotation through a Hilbert space, statistical uncertainty, and related ideas are not unique to quantum mechanics (which is unfortunately too often a consequence of the way quantum mechanics is typically introduced in undergraduate curricula, at least in the US). We also came to a consensus that in each course, there should be clear explanations of what prototypical systems are analytically solvable, what prototypical systems are not analytically solvable, and why (in each case).

2023-11-01

Contravariant and Covariant Objects in Matrix Notation

For many years when and since I was in college, I wondered whether it might be possible to consistently represent contravariant & covariant objects using vector & matrix notation. In particular, when I learned about the idea of covariant representations of [invariant] vectors being duals to contravariant representations of [invariant] vectors, meaning that if a contravariant representation of a [invariant] vector can be seen as a column vector, then a covariant representation of a [invariant] vector can be seen as a row vector, I wondered how it would be possible to represent the fully covariant metric tensor as a metric tensor if it multiplies a contravariant representation of a [invariant] vector (i.e. a column vector) to yield a covariant representation of a [invariant] vector (i.e. a row vector), especially as traditionally in linear algebra, a matrix acting on a column vector yields another column vector (while transposition, though linear in the sense of respecting addition and scalar multiplication, cannot be represented simply as the action of another matrix). At various points, I've wondered if this means that fully contravariant or fully covariant representations of multi-index tensors should be represented as columns of columns or rows of rows, and I've tried to play around with these ideas more. This post is not the first to explore such ideas even online, as I came across notes online by Viktor T. Toth [LINK], but this post is my attempt to flesh out these ideas further. Follow the jump to see more. Throughout this post, I will work with the notation of 2 spatial indices, in which the fully covariant representation of the metric tensor \( g_{ij} = \vec{e}_{i} \cdot \vec{e}_{j} \) might not be Euclidean, where indices will use English letters \( i, j, k, \ldots \in \{1, 2\} \), where superscripts do not imply exponents, and where multiple superscripts do not imply single numbers (for example, \( g_{12} \) is the fully covariant component of the metric tensor with first index 1 and second index 2, not the covariant component at index 12 of a single-index tensor (vector)); extensions to spacetime (where the convention is to use indices labeled by Greek letters) and in particular to 3 spatial + 1 temporal dimensions are trivial. Additionally, Einstein summation will be assumed, and all tensors (including vectors & scalars) are assumed to be real-valued. Finally, I will do my best to ensure that when indices are raised or lowered, the ordering of indices is clear (as examples, distinguishing \( T^{i}_{\, j} \) from \( T_{i}^{\, j} \) instead of ambiguously using \( T^{i}_{j} \) or \( T^{j}_{i} \)), but this will depend on the quality of LaTeX rendering in this post.

2023-06-11

Book Review: "How Not to Be Wrong" by Jordan Ellenberg

I've recently read the book How Not to Be Wrong by Jordan Ellenberg. As the author states in the introduction, it is an exposition of simple yet profound ideas in mathematics, meant for laypeople. Topics include nonlinear phenomena (in opposition to naïve linear extrapolation), probability, Bayesian reasoning, and statistical testing of hypotheses. All chapters refer to many examples in politics, economics, and everyday life to make the concepts easier for laypeople to digest.

I found the book to be fairly easy to follow. I can't say that I learned much in terms of concepts, as these are all concepts that I've come across one way or another in school, college, graduate school, or my work now, though I did appreciate the discussion of how conspiracy theorists like to add hypotheses after the fact to make a conspiracy theory harder to fully disprove, how the fact that random fluctuations in many phenomena observed over time are time-reversal invariant implies that the phenomenon of regression toward the mean is also time-reversal invariant in a probabilistic sense, and the intuitive explanations of common causes & common effects in leading to correlations between random variables that are otherwise not causally connected. Additionally, I felt like this book did a better job than the book Algorithms to Live By by Brian Christian & Tom Griffiths (which I have reviewed on this blog before [LINK]) in having some structure in the progression from one chapter to the next and in using topics from earlier chapters in later chapters even though this book, unlike that book, didn't pretend to have a unified message. My only quibbles are the claim that the impossibility of accurately running the fundamental equations describing atmospheric & oceanic dynamics for more than 2 weeks implies impossibility in forecasting through other methods (like machine learning models looking for patterns in weather effects & progression) and the fact that the chapter connecting ideas from probability, geometry, and signal processing (particularly around error correction) took me a fair bit of effort to follow (unlike the other chapters, which tells me that laypeople will likely struggle with that chapter much more). Additionally, I think readers should be aware that the author often makes reference to sports that are mostly popular in the US and to US politics and that the author at a few points espouses more liberal or progressive political views (though I think such espousal is not gratuitous but is done in a way that fits well with broader discussions of assumptions underlying mathematical, political, and legal judgments). Overall, I think the author has done a good job of fulfilling the goal of communicating these ideas to a lay audience, so I recommend this book to anyone who might be interested in these ideas.

2022-12-02

Fundamental Theorem of Calculus for Functionals

I happened to think more about the idea of recovering a functional by somehow integrating its functional derivative. In the process, I realized that certain ideas that I would have to consider make this post a natural follow-up to a recent post [LINK] about mapping scalars to functions. This will become clear later in this post.

For a single variable, a function \( f(x) \) has an antiderivative \( F(x) \) such that \( f(x) = \frac{\mathrm{d}F}{\mathrm{d}x} \). One statement of the fundamental theorem of calculus is that this implies that \[ \int_{a}^{b} f(x)~\mathrm{d}x = F(b) - F(a) \] for these functions. In turn, this means \( F(x) \) can be extracted directly from \( f(x) \) through \[ F(x) = \int_{x_{0}}^{x} f(x')~\mathrm{d}x' \] in which \( x_{0} \) is chosen such that \( F(x_{0}) = 0 \).

For multiple variables, a conservative vector field \( \mathbf{f}(\mathbf{x}) \) in which \( \mathbf{f} \) must have the same number of components as \( \mathbf{x} \) can be said to have a scalar antiderivative \( F(\mathbf{x}) \) in the sense that \( \mathbf{f} \) is the gradient of \( F \), meaning \( \mathbf{f}(\mathbf{x}) = \nabla F(\mathbf{x}) \); more precisely, \( f_{i}(x_{1}, x_{2}, \ldots, x_{N}) = \frac{\partial F}{\partial x_{i}} \) for all \( i \in \{1, 2, \ldots, N \} \). (Note that if \( \mathbf{f} \) is not conservative, then it by definition cannot be written as the gradient of a scalar function! This is an important point to which I will return later in this post.) In such a case, a line integral (which, as I will emphasize again later in this post, is distinct from a functional path integral) from vector point \( \mathbf{a} \) to vector point \( \mathbf{b} \) of \( \mathbf{f} \) can be computed as \( \int \mathbf{f}(\mathbf{x}) \cdot \mathrm{d}\mathbf{x} = F(\mathbf{b}) - F(\mathbf{a}) \); more precisely, this equality holds along any contour, so if a contour is defined as \( \mathbf{x}(s) \) for \( s \in [0, 1] \), no matter what \( \mathbf{x}(s) \) actually is, as long as \( \mathbf{x}(0) = \mathbf{a} \) and \( \mathbf{x}(1) = \mathbf{b} \) hold, then \[ \sum_{i = 1}^{N} \int_{0}^{1} f_{i}(x_{1}(s), x_{2}(s), \ldots, x_{N}(s)) \frac{\mathrm{d}x_{i}}{\mathrm{d}s} \mathrm{d}s = F(\mathbf{b}) - F(\mathbf{a}) \] must also hold. This therefore suggests that \( F(\mathbf{x}) \) can be extracted from \( \mathbf{f}(\mathbf{x}) \) by relabeling \( \mathbf{x}(s) \to \mathbf{x}'(s) \), \( \mathbf{a} \) to a point such that \( F(\mathbf{a}) = 0 \), and \( \mathbf{b} \to \mathbf{x} \). Once again, if \( \mathbf{f}(\mathbf{x}) \) is not conservative, then it cannot be written as the gradient of a scalar field \( F \), and the integral \( \sum_{i = 1}^{N} \int_{0}^{1} f_{i}(x_{1}(s), x_{2}(s), \ldots, x_{N}(s)) \frac{\mathrm{d}x_{i}}{\mathrm{d}s} \mathrm{d}s \) will depend on the specific choice of \( \mathbf{x}(s) \), not just the endpoints \( \mathbf{a} \) and \( \mathbf{b} \).

For continuous functions, the generalization of a vector \( \mathbf{x} \), or more precisely \( x_{i} \) for \( i \in \{1, 2, \ldots, N\} \), is a function \( x(t) \) where \( t \) is a continuous dummy index or parameter analogous to the discrete index \( i \). This means the generalization of a scalar field \( F(\mathbf{x}) \) is the scalar functional \( F[x] \). What is the generalization of a vector field \( \mathbf{f}(\mathbf{x}) \)? To be precise, a vector field is a collection of functions \( f_{i}(x_{1}, x_{2}, \ldots, x_{N}) \) for all \( i \in \{1, 2, \ldots, N \} \). This suggests that its generalization should be a function of \( t \) and must somehow depend on \( x(t) \) as well. It is tempting therefore to write this as \( f(t, x(t)) \) for all \( t \). However, although this is a valid subset of the generalization, it is not the whole generalization, because vector fields of the form \( f_{i}(x_{i}) \) are collections of single-variable functions that do not fully capture all vector fields of the form \( f_{i}(x_{1}, x_{2}, \ldots, x_{N}) \) for all \( i \in \{1, 2, \ldots, N \} \). As a specific example, for \( N = 2 \), the vector field with components \( f_{1}(x_{1}, x_{2}) = (x_{1} - x_{2})^{2} \) and \( f_{2}(x_{1}, x_{2}) = (x_{1} + x_{2})^{3} \) cannot be written as just \( f_{1}(x_{1}) \) and \( f_{2}(x_{2}) \), as \( f_{1} \) depends on \( x_{2} \) and \( f_{2} \) depends on \( x_{1} \) as well. Similarly, in the generalization, one could imagine a function of the form \( f = \frac{x(t)}{x(t - t_{0})} \mathrm{exp}(-(t - t_{0})^{2}) \); in this case, it is not correct to write it as \( f(t, x(t)) \) because the dependence of \( f \) on \( x \) at a given dummy index value \( t \) comes through not only \( x(t) \) but also \( x(t - t_{0}) \) for some fixed parameter \( t_{0} \). Additionally, the function may depend not only on \( x \) per se but also on derivatives \( \frac{\mathrm{d}^{n} x}{\mathrm{d}t^{n}} \); the case of the first derivative \( \frac{\mathrm{d}x}{\mathrm{d}t} = \lim_{t_{0} \to 0} \frac{x(t) - x(t - t_{0})}{t_{0}} \) illustrates the connection to the aforementioned example. Therefore, the most generic way to write such a function is effectively as a functional \( f[x; t] \) with a dummy index \( t \). The example \( f = \frac{x(t)}{x(t - t_{0})} \mathrm{exp}(-(t - t_{0})^{2}) \) can be formalized as \( f[t, x] = \int_{-\infty}^{\infty} \frac{x(t')}{x(t' - t_{0})} \mathrm{exp}(-(t' - t_{0})^{2}) \delta(t - t')~\mathrm{d}t' \) where the dummy index \( t' \) is the integration variable while the dummy index \( t \) is free. (For \( N = 3 \), the condition of a vector field being conservative is often written as \( \nabla \times \mathbf{f}(\mathbf{x}) = 0 \). I have not used that condition in this post because the curl operator does not easily generalize to \( N \neq 3 \).)

If a functional \( f[x; t] \) is conservative, then there exists a functional \( F[x] \) (with no free dummy index) such that \( f \) is the functional derivative \( f[x; t] = \frac{\delta F}{\delta x(t)} \). Comparing the notation between scalar fields and functionals, \( \sum_{i} A_{i} \to \int A(t)~\mathrm{d}t \) and \( \mathrm{d}x_{i} \to \delta x(t) \), in which \( \delta x(t) \) is a small variation in a function \( x \) specifically at the index value \( t \) and nowhere else. This suggests a generalization of the fundamental theorem of calculus to functionals as follows. If \( a(t) \) and \( b(t) \) are fixed functions, then \( \int_{-\infty}^{\infty} \int f[x; t]~\delta x(t)~\mathrm{d}t = F[b] - F[a] \). More precisely, a path from the function \( a(t) \) to the function \( b(t) \) at every index value \( t \) can be parameterized by \( s \in [0, 1] \) by the map \( s \to x(t, s) \) which is a function of \( t \) for each \( s \) such that \( x(t, 0) = a(t) \) and \( x(t, 1) = b(t) \); this is why I linked this post to the most recent post on this blog. With this in mind, the fundamental theorem of calculus becomes \[ \int_{-\infty}^{\infty} \int_{0}^{1} f[x(s); t] \frac{\partial x}{\partial s}~\mathrm{d}s~\mathrm{d}t = F[b] - F[a] \] where, in the integrand, the argument \( x \) in \( f \) has the parameter \( s \) explicit but the dummy index \( t \) implicit; the point is that this equality holds regardless of the specific parameterization \( x(t, s) \) as long as \( x \) at the endpoints of \( s \) satisfies \( x(t, 0) = a(t) \) and \( x(t, 1) = b(t) \). This also means that \( F[x] \) can be recovered if \( b(t) = x(t) \) and \( a(t) \) is chosen such that \( F[a] = 0 \), in which case \[ F[x] = \int_{-\infty}^{\infty} \int_{0}^{1} f[x'(s); t]~\frac{\partial x'}{\partial s}~\mathrm{d}s~\mathrm{d}t \] (where \( x(t, s) \) has been renamed to \( x'(t, s) \) to avoid confusion with \( x(t) \)). If \( f[x; t] \) is not conservative, then there is no functional \( F[x] \) whose functional derivative with respect to \( x(t) \) would yield \( f[x; t] \); in that case, with \( x(t, 0) = a(t) \) and \( x(t, 1) = b(t) \), the integral \( \int_{-\infty}^{\infty} \int_{0}^{1} f[x(s); t] \frac{\partial x}{\partial s}~\mathrm{d}s~\mathrm{d}t \) does depend on the specific choice of parameterization \( x(t, s) \) with respect to \( s \) and not just on the functions \( a(t) \) and \( b(t) \) at the endpoints of \( s \).

As an example, consider from a previous post [LINK] the nonrelativistic Newtonian action \[ S[x] = \int_{-\infty}^{\infty} \left(\frac{m}{2} \left(\frac{\mathrm{d}x}{\mathrm{d}t}\right)^{2} + F_{0} x(t) \right)~\mathrm{d}t \] for a particle under the influence of a uniform force \( F_{0} \) (which may vanish). The first functional derivative is \[ f[x; t] = \frac{\delta S}{\delta x(t)} = F_{0} - m\frac{\mathrm{d}^{2} x}{\mathrm{d}t^{2}} \] and its vanishing would yield the usual equation of motion. The action itself vanishes for \( x(t) = 0 \), which will be helpful when using the fundamental theorem of calculus to recover the action from the equation of motion. In particular, one can parameterize \( x'(t, s) = sx(t) \) such that \( x'(t, 0) = 0 \) and \( x'(t, 1) = x(t) \). This gives the integral \( \int_{0}^{1} \left(F_{0} - ms\frac{\mathrm{d}^{2} x}{\mathrm{d}t^{2}}\right)x(t)~\mathrm{d}s = F_{0} x(t) - \frac{m}{2} x(t) \frac{\mathrm{d}^{2} x}{\mathrm{d}t^{2}} \). This is then integrated over all \( t \), so the first term is identical to the corresponding term in the definition of \( S[x] \), and the second term becomes the same as the corresponding term in the definition of \( S[x] \) after integrating over \( t \) by parts and setting the boundary conditions that \( x(t) \to 0 \) for \( |t| \to \infty \). (Other boundary conditions may require more care.) In any case, the parameterization \( x'(t, s) = sx(t) \) is not the only choice that could fulfill the boundary conditions; the salient point is that any parameterization fulfilling the boundary conditions would yield the correct action \( S[x] \).

I considered that example because I wondered whether any special formulas need to be considered if \( f[x; t] \) depends explicitly on first or second derivatives of \( x(t) \), as might be the case in nonrelativistic Newtonian mechanics. That example shows that no special formulas are needed because even if the Lagrangian explicitly depends on the velocity \( \frac{\mathrm{d}x}{\mathrm{d}t} \), the action \( S \) only explicitly depends as a functional on \( x(t) \), so proper application of functional differentiation and regular integration by parts will ensure proper accounting of each piece.

This post has been about the fundamental theorem of calculus saying that the 1-dimensional integral of a function in \( N \) dimensions along a contour, if that function is conservative, is equal to the difference between the two endpoints of its scalar antiderivative. This generalizes easily to infinite dimensions and continuous functions instead of finite-dimensional vectors. There is another fundamental theorem of calculus saying that the \( N \)-dimensional integral in a finite volume of the scalar divergence of an \( N \)-dimensional vector function, if that volume has a closed orientable surface, is equal to the \( N - 1 \)-dimensional integral of the inner product of that function with the normal vector (of unit 2-norm) at every point on the surface across the whole surface, meaning \[ \int_{V} \sum_{i = 1}^{N} \frac{\partial f_{i}}{\partial x_{i}}~\mathrm{d}V = \oint_{\partial V} \sum_{i = 1}^{N} f_{i}(x_{1}, x_{2}, \ldots, x_{N}) n_{i}(x_{1}, x_{2}, \ldots, x_{N})~\mathrm{d}S \] where \( \sum_{i = 1}^{N} |n_{i}(x_{1}, x_{2}, \ldots, x_{N})|^{2} = 1 \) for every \( \mathbf{x} \). From a purely formal perspective, this could generalize to something like \( \int_{V} \int_{-\infty}^{\infty} \frac{\delta f[x; t]}{\delta x(t)}~\mathrm{d}t~\mathcal{D}x = \oint_{\partial V} \int_{-\infty}^{\infty} f[x; t]n[x; t]~\mathrm{d}t~\mathcal{D}x \) having generalized \( \frac{\partial}{\partial x_{i}} \to \frac{\delta}{\delta x(t)} \), \( \prod_{i} \mathrm{d}x_{i} \to \mathcal{D}x \), and \( n_{i}(\mathbf{x}) \to n[x; t] \) where \( n[x; t] \) is normalized such that \( \int_{-\infty}^{\infty} |n[x; t]|^{2}~\mathrm{d}t = 1 \) for all \( x(t) \) on the surface. However, this formalism may be hard to further develop because the space has infinite dimensions. Even when working in a countable basis, it might not be possible to characterize an orientable surface enclosing a volume in an infinite-dimensional space; the surface is also infinite-dimensional. While the choice of basis is arbitrary, things become even less intuitive when choosing to work in an uncountable basis.

2022-11-01

Mapping Scalars to Functions

In just over a year, I've written three posts for this blog about functionals, specifically about their application to probability theory [LINK], finding their stationary points [LINK], and the use of their stationary points in classical mechanics [LINK]. As a reminder, a functional is an object that maps a space of functions to a space of numbers. This got me thinking about what the reverse, namely an object that maps a space of numbers to a space of functions, looks like. To be clear, this is not the same as an ordinary function which, as an element in a space of functions, maps a space of numbers to a space of numbers.

As I thought about it more, I realized that this is a bit easier to understand and therefore more commonly encountered than a functional. An extremely glib way to describe such an object is a function of multiple variables. However, it may be more enlightening to describe this in further detail to avoid potentially deceptive images that may arise from that glib description.

In the discrete case, the matrix elements \( A_{ij} \) can be described as a map from integers to vectors, in which an integer \( j \) is associated with a vector whose elements indexed by an integer \( i \) are \( A_{ij} \). This is the essential idea behind seeing the columns of the matrix with elements \( A_{ij} \) as a collection of vectors. Formally, this maps \( i \to (j \to A_{ij}) \) where the map \( j \to A_{ij} \) defines a vector indexed by the free variable \( i \).

Similarly, in the continuous case, the function elements \( f(x, y) \) can be described as a map from numbers to functions, in which a number \( y \) is associated with a function whose elements indexed by a number \( x \) are \( f(x, y) \). Formally, this maps \( x \to (y \to f(x, y)) \) where the map \( y \to f(x, y) \) defines a function indexed by the free variable \( x \). These ideas are foundational to the development of more abstract notions of functions, like lambda calculus.

2022-06-01

Nonlocality and Infinite LDOS in Lossy Media

While I have written many posts on this blog about various topics in physics or math unrelated to my graduate work as well as posts promoting papers from my graduate work, it is rare that I've written direct technical posts about my graduate work. It is even more unusual that I should be doing so 2 years after leaving physics as a career. However, I felt compelled to do so after meeting again with my PhD advisor (a day before the Princeton University 2020 Commencement, which was held in person after a delay of 2 years due to this pandemic), as we had a conversation about the problem of infinite local density of states (LDOS) in a lossy medium.

Essentially, the idea is the following. Working in the frequency domain, the electric field produced by a polarization density in any EM environment is \( E_{i}(\omega, \vec{x}) = \int G_{ij}(\omega, \vec{x}, \vec{x}')P_{j}(\omega, \vec{x}')~\mathrm{d}^{3} x' \) which can be written in bra-ket notation (dispensing with the explicit dependence on frequency) as \( |\vec{E}\rangle = \hat{G}|\vec{P}\rangle \). The LDOS is proportional to the power radiated by a point dipole and can be written as \( \mathrm{LDOS}(\omega, \vec{x}) \propto \sum_{i} \mathrm{Im}(G_{ii}(\omega, \vec{x}, \vec{x})) \). This power should be finite as long as the power put into the dipole to keep it oscillating forever at a given frequency \( \omega \) is finite. However, there appears to be a paradox in that if the position \( \vec{x} \) corresponds to a point embedded in a local lossy medium, the LDOS diverges there.

I wondered if an intuitive explanation could be that loss should properly imply the existence of energy leaving the system by traveling out of its boundaries, so the idea of a medium that is local everywhere (in the sense that the susceptibility operator takes the form \( \chi_{ij}(\omega, \vec{x}, \vec{x}') = \chi_{ij}(\omega, \vec{x})\delta^{3} (\vec{x} - \vec{x}') \) at all positions) and is lossy at every point in its domain may not be well-posed as energy is somehow disappearing "into" the system instead of leaving it. Then, I wondered if the problem may actually be with locality and whether a nonlocal description of the susceptibility could help. This is where my graduate work could come in. Follow the jump to see a very technical sketch of how this might work (as I won't work out all of the details myself).

2022-05-01

Book Review: "Algorithms to Live By" by Brian Christian & Tom Griffiths

I've recently read the book Algorithms to Live By by Brian Christian & Tom Griffiths. This book shows how many problems & heuristics in computer science can be applied to explain or improve human decision-making. Each chapter focuses on a certain class of problems or issues. Such classes include the optimal stopping problem, the multi-armed bandit problem, searching & sorting, task scheduling, Bayesian inference, overfitting data, constraint relaxation, random stimulus, communication protocols, and social interaction. Additionally, most chapters try to show how results from computer science can either improve or justify certain human behaviors.

This book was frustrating for me to read. If it had fully met my expectation that it would show, in a unified & consistent way, how these computer science problems apply to human behavior and connect to each other, I would be singing its praises. If it had completely failed, I'd be happy to rhetorically trash this book. Instead, I found that each chapter would be a great vignette on its own, and each chapter showed the great potential of what the book could have been, but the book failed to live up to that potential. First, there was very little connection among the chapters, and any acknowledgment that the authors did make of such connections was almost always superficial instead of deeply insightful. For example, the respective chapters about the optimal stopping problem, caches, and overfitting each could have been so much better with greater discussion about the connection to social pressure & game theory, yet those topics were discussed only in the last chapter, which I think was a mistake. Second, only in the concluding section did the authors make clear that they wanted to either improve or justify human behavior with each class of problems or issues. This because clear over the course of reading the book, yet there was very little guidance in each chapter about whether improvement versus justification would be the goal. Perhaps the worst offender was the chapter about constraint relaxation, as there was little connection to human behavior in a way that would be obvious to lay readers. These problems meant that reading the last numbered chapter (about game theory) and the conclusion felt simultaneously wonderful for finally seeing these concepts discussed clearly and maddening for knowing that the book could have been so much better if these ideas had been more consistently executed through the book.

There are two other minor criticisms I have of the book too. First, the chapter about overfitting seems to use the word "overfitting" to mean too many different things, which is ironic and undermines any clarity that the discussion could have provided. Second, the chapter about randomized algorithms attempts to make a tenuous connection between randomized algorithms used in computer science and the way that random mental stimuli can produce very creative responses in people, but it never makes clear whether the latter result is true at an individual level or only holds statistically for large populations.

Overall, I think the author's goals were laudable and that each chapter is interesting to read in isolation. However, other readers may be disappointed, as I was, in the way that the authors fail to synthesize many of the ideas across chapters in a smooth & unified manner. Thus, I would advise that readers who may be interested in these topics go into this book with lower expectations.

2022-04-04

FOLLOW-UP: How to Tell Whether a Functional is Extremized

This post is a follow-up to an earlier post (link here) about how to tell whether a stationary point of a functional is a maximum, minimum, or saddle point. In particular, as I thought about it more, I realized that using the analogy to discrete vectors could help when formulating a more general expression for the second derivative of the nonrelativistic classical action for a single degree of freedom (i.e. the corresponding Hessian operator). Additionally, I thought of a few other examples of actions whose Hessian operators are positive-definite. Finally, I've thought more about how to express these equations for systems with multiple degrees of freedom (DOFs) as well as for fields and about how these ideas connect to the path integral formulation of quantum mechanics. Follow the jump to see more

2022-03-05

How to Tell Whether a Functional is Extremized

I happened to be thinking recently about how to tell when a functional is extremized. Examples in physics include minimizing the ground state energy of an electronic system expressed as an approximate density functional \( E[\rho] \) with respect to the electron density \( \rho \) or maximizing the relativistic proper time \( \tau \) of a classical particle with respect to a path through spacetime. Additionally, finding the points of stationary action that lead to the Euler-Lagrange equations of motion is often called "minimization of the action", but I can't recall ever having seen a proof that the action is truly minimized (as opposed to reaching a saddle point). This got me to think more about the conditions under which a functional is truly maximized or minimized as opposed to reaching a saddle point. Follow the jump to see more. I will frequently refer to concepts presented in a recent post (link here), including the relationships between functionals of vectors & functionals of functions. Additionally, for simplicity, all variables and functions will be real-valued.

2021-10-04

Functionals in Probability and Bayesian Inference

My work on transportation policy research in part involves conducting & analyzing surveys of people's travel behaviors & attitudes. Analyzing survey data requires an understanding of basic probability and statistics, which is an area that I previously felt I had just enough knowledge of to get by when learning about statistical physics but that I need to build more practical skills in now. In the process of refreshing my understanding of probability and statistics, I thought more about Bayes's theorem. In the context of hypothesis testing or inference, Bayes's theorem can be stated as follows: given a hypothesis \( \mathrm{H} \) and data \( \mathrm{D} \) such that the likelihood of measuring the data under that hypothesis is \( \operatorname{P}(\mathrm{D}|\mathrm{H}) \), and given a prior probability \( \operatorname{P}(\mathrm{H}) \) associated with that hypothesis, the posterior probability of that hypothesis is \[ \operatorname{P}(\mathrm{H}|\mathrm{D}) = \frac{\operatorname{P}(\mathrm{D}|\mathrm{H})\operatorname{P}(\mathrm{H})}{\operatorname{P}(\mathrm{D})} \] given the data. The key is that the denominator is evaluated as a sum \[ \operatorname{P}(\mathrm{D}) = \sum_{\mathrm{H}'} \operatorname{P}(\mathrm{D}|\mathrm{H}')\operatorname{P}(\mathrm{H}') \] where the label \( \mathrm{H}' \) runs over all possible hypothesis.

In practice, however, the set of hypotheses doesn't literally encompass all hypotheses but encompasses only one particular type of function with one or a few free parameters which then go into the prior probability distribution. For one free parameter, if the hypothesis specifies only the value of the (assumed continuous) parameter \( \theta \), if the prior probability of that hypothesis is given by a density \( f_{\mathrm{H}}(\theta)\), and the likelihood of measuring a (assumed continuous) data vector \( D \) under that hypothesis is the density \( f(D|\theta) \), then Bayes's theorem gives \[ f_{\mathrm{H}}(\theta|D) = \frac{f(D|\theta) f_{\mathrm{H}}(\theta)}{\int f(D|\theta') f_{\mathrm{H}}(\theta')~\mathrm{d}\theta'} \] as the posterior probability density under that hypothesis given the data.

I understood that in most cases, a single class of likelihood functions varied through a single parameter is good enough, and especially for the purposes of pedagogy, it is useful to keep things simple & analytical. Even so, I was more broadly unsatisfied with the lack of explanation for how to more generally consider summing over all possible hypotheses.

This post is my attempt to address some of those issues. Follow the jump to see more explanation as well as discussion of other tangentially related philosophical points.

2020-10-19

Book Review: "The Drunkard's Walk" by Leonard Mlodinow

I've recently reread the book The Drunkard's Walk by Leonard Mlodinow, which is a book about many ways that probabilistic phenomena occur in daily life and what the consequences are for understanding individual & collective decisions. I say "reread" because the first time I read it was in high school (as I recall, although I don't remember exactly when, though I did mention it in a review for a different book, saying then that I didn't finish it because I didn't find it as engaging as the book in that review); a few days ago, I happened to see it on top a stack of books, and I figured it would be nice to reread for a few reasons. First, I have learned a lot of science, and my worldview has developed & matured a lot, since I was in high school, so I thought it would be good to see how this book would hold up in my view in that context. Second, I figured it would be nice to read a book about probabilistic phenomena, as it wasn't something that I had to worry much about in my college studies or in my PhD work (which is a little ironic, given that van der Waals forces and radiative heat transfer are phenomena of statistical thermodynamic fluctuations, but it turns out that certain mathematical formulations hide all essential randomness), it will be relevant to my postdoctoral work as I get more into travel surveys with associated statistical analysis, concepts like base rate fallacies are relevant for things like false positive result rates for tests associated with this coronavirus (please note that I am not a public health expert, and please consult governmental public health agencies for guidance with respect to this ongoing pandemic), and I've been thinking over the last several months about how many of the conceptual quandaries associated with quantum mechanics can actually be tied to questions of whether probability is emergent versus fundamental.

The book is not too long, and it is a quick & engaging read. The author uses many interesting examples to motivate the discussion of fallacious reasoning in the context of probability as well as ways that probability enters daily life even in areas where people expect more determinism. There are also many interesting historical anecdotes about the development of probability theory, especially how ancient Greece and certain medieval European societies believed that any discussion of uncertainty would go against their conceptions of a pure & deterministic universe (whatever the prime mover might be). Also, in the tenth chapter, there is an interesting discussion of how the development of chaos theory itself is an example of the unpredictable & seemingly random nature of human life (though I didn't like the conflation of chaos theory itself with probability, as chaos is a separate mathematical phenomena that can emerge in purely deterministic systems). Additionally, in the tenth chapter, I appreciated how the author is careful to state that determinism is a bad model only of human behavior (at individual & societal levels) and makes no claim about the applicability of determinism to the universe at large, and how the author makes a call for humility and for rewarding people based on their character instead of perpetuating beliefs that people who are successful are wholly responsible for their successes while people who are in marginalized circumstances are somehow rightfully being punished for past mistakes. Overall, I think the book does a good job of achieving its purpose of illustrating to lay readers how ubiquitous probabilistic phenomena are in even seemingly deterministic aspects of daily life.

Before getting into other criticisms, I should point out that my copy of this book has several printing errors (mostly missing words) and a few typographical errors, but these occurred maybe once every 10 pages (based on an instinctive guess), so these therefore didn't affect my understanding of the book. Also, the author errs in claiming that Germanic rule in the Dark Ages (commonly understood to be the medieval period) preceded the ancient Roman civilization, but this again doesn't undercut the overall argument.

Where this book falls short is in fulfilling its purpose of diving deeper into the implications of such randomness for human behavior at individual and societal levels; the author's sloppy treatment of human behavior is a recurring problem throughout the book. In particular, there are a few related broad issues that come up at various points through the book. The first is the question of how to reconcile the apparent randomness of daily events (including the phenomenon of regression to the mean) believed to be deterministic with the real phenomena of collective self-fulfilling prophecies (including emergent segregation of social groups to reinforce outcomes that are believed to be deterministic even if they are not, thereby reinforcing determinism in itself). The second is the treatment of things like superstitions as examples of self-fulfilling prophecies, even if the superstitions have no effects in their contents but may change mindsets enough to change outcomes. The author doesn't do a good job of addressing many of these issues throughout the book, and only partially acknowledges the power of self-fulfilling prophecies at the end of the book (in the tenth chapter); the author makes it seem like a slow & methodical build-up to a satisfying conclusion, but frankly, the discussion of these issues could have been a lot more clear & concise and could have come much sooner in the book. The discussion of superstition in particular is rife with condescension, as the author never acknowledges how superstitions may change mindsets & lead to self-fulfilling prophecies but instead summarily dismisses them as silly relics mostly of a bygone era, reinforced by statements about how science and religion were irreparably separated with the trial of Galileo with no nuanced discussion of how religious beliefs (even if not organized religious institutions per se, to the extent that was the case before Galileo) played a role in motivating scientific discoveries even after Galileo. Another example is how the author glibly dismisses many claims of clusters of environmentally-caused cancer; it may well be true that some cases are due to biased statistical analysis after the fact, but it doesn't really address why inequitable outcomes seem to occur so frequently in this context, and it doesn't do justice to the gravity of the problem. (UPDATE: I recognize that my argument against the author's treatment of the incidence of environmentally-caused cancer can easily be dismissed as an overly emotional reaction that is not justified by the statistics, so it is worth clarifying that further. My concern is that the statistical arguments that claim that environmentally-caused cancer is not really a problem, and that those who claim it is a problem only do so by drawing arbitrary boundaries after the fact to inflate apparent concentrations of carcinogens in specific areas, may themselves be riven with the same sorts of bias that are perpetuated in situations like machine learning determining the provision of health care, but are cast in a way that seems "neutral" and therefore "superior" to "emotionally-driven" arguments.) Furthermore, although there is discussion of both the failures of superficial statistical arguments in favor of DNA testing in the criminal justice system and of the way that Bayesian analysis can systematically codify learning of new information in terms of probabilities, there is little discussion of how these issues can combine in toxic ways to perpetuate existing societal biases under the veneer of formal Bayesian analysis (as occurs with machine learning now); I admit that I wouldn't have been thinking about this as much had I not read and reviewed the book Weapons of Math Destruction by Cathy O'Neil, but similar examples were already available at the time that the book that I review in this post was being written. I really see an essential condescension and a lack of humility throughout the book in these discussions of human behavior, masked by the pithy & irreverent writing, that are at odds with the author's own calls for humility & deeper understanding.

There are other aspects of human behavior that this book fails to adequately capture; these may technically be beyond the scope of this book, but I think they are worth noting anyway, as they speak to larger problems with the ability of people (even those well-trained in STEM fields) to really understand probability theory. The second chapter goes over many examples of how, in the technical language of probability, given events \( A \) and \( B \), certain questions can be framed such that laypeople and professional specialists (particularly doctors & lawyers) fall into the trap of believing that \( \operatorname{Pr}(A \cap B) > \operatorname{Pr}(A) \) even though the opposite is mathematically always true. However, I can already see that the phrasing of many of those questions, particularly the way that events \( A \) & \( B \) are juxtaposed (especially if \( B \) is additional information that may be relevant to the assessment of \( A \)), may make people believe either that what should be interpreted as \( \operatorname{Pr}(A) \) is actually \( \operatorname{Pr}(A \cap \neg B) \), in which case \( \operatorname{Pr}(A \cap B) > \operatorname{Pr}(A \cap \neg B) \) could in fact be true, or that what should be interpreted as \( \operatorname{Pr}(A \cap B) \) is actually the conditional probability \( \operatorname{Pr}(A|B) \), in which case \( \operatorname{Pr}(A|B) > \operatorname{Pr}(A) \) could in fact be true. This speaks more to the way that natural human language is unsuited to the subtleties of the language of probability theory, yet rather than address these possibilities, the author again leaves the discussion there, implying disdain for people who are too stupid to know better. Another problem is that throughout the book, the author raises the question of how to determine whether a particular sequence of observations of outcomes for a process that may be random reflects a specific probability distribution model, but never clearly explains how to do this in practice, instead only giving hints about this through various examples. This is related to the question of why one may prefer an explanation based on probabilities than based on deterministic phenomena, particularly for small sample sizes. For this, I will give an example. Consider exactly 5 observations of an event, which has binary outcomes (either success or failure), for which no other observations are made, and for which in all of those 5 observations, success occurs every single time. Intuitively, laypeople might be inclined to believe that there is a deterministic cause of this, while if a probability theorist were to initially believe that this is consistent with a binomial distribution with \( (N, p) = (5, 0.6) \) but then later revise this to \( (N, p) = (5, 0.99) \), laypeople could reasonably wonder why this would be justified, and why the probability theorist refuses to believe in the possibility of some deterministic causal relationship. Of course, this is a contrived example, I understand why causation needs to be proved as an alternative to a null hypothesis, and I understand that probability distributions closer to uniform probabilities are favored as those that maximize entropy (which essentially means that subject to certain known constraints, the probability distribution that best reflects the state of ignorance about a system is closest to uniform), but the author does not properly explain these points. Finally, the broadest problem with this book is that the author only superficially acknowledges the issue that if every calculation in probability theory or statistics, whether of a certain event happening, a string of events being a true "hot streak", or a model fitting data correctly, is itself a probability, then the aforementioned disconnect of this language of probability from natural human language makes it difficult to translate probabilities into robust rules for deterministic (usually binary) decisions that laypeople must make; this is related to the idea that in game theory, a single person playing a single-shot game cannot play a mixed strategy, and the concept of a mixed strategy only makes sense in the context of observing a large ensemble of independent players, possibly playing repeatedly. Perhaps asking the author to address this problem is too much, but I still feel like such failures diminish the book in comparison to its hype.

Without hyping my own credentials, I admit that it is possible that my reaction to this book is more of a reflection of my greater experience with STEM, humanities, and social science fields and with science education/communication compared to when I was in high school. Furthermore, just as this book exhorts, I cannot be overcome by either positivity bias or negativity bias; it would only be fair to take the good & bad parts of this book together as appropriate, without believing that one outdoes or cancels the other. Given this, I can't really make a strong recommendation that readers should or should not read this book.

2020-07-01

Classical Phase Space Densities for One or a Few Particles

This is the first time in several years that I've done a post about physics that didn't have to do with my research. This came about from thinking about applying techniques in statistical physics to game theory; although I still have a lot more to learn about that and need to do more to flesh out those ideas, it occurred to me in the process that I never had such a good intuition for the phase space density in classical mechanics, and notes that I've found online focus almost exclusively on the phase space density of a large number of particles in an explicitly statistical treatment. I intend to use this post to shed light on why this may be the case, help build intuition for how things like the Liouville equation work for simple systems of one or a few particles, and reinforce the notion that there is no classical analogue to the phenomenon of a multi-particle entangled quantum state yielding a mixed single-particle state under a partial trace. Follow the jump to see more.

2017-12-18

Book Review: "Hidden Figures" by Margot Lee Shetterly

I've recently read the book Hidden Figures by Margot Lee Shetterly. It weaves together the true stories of a few particular mathematicians, who happened to be black women (among a larger group of such female black mathematicians), who made extremely important contributions to the development of American warplanes in WWII and then spacecrafts during the 1950s and 1960s, including the crafts that took John Glenn to space and then the Apollo 11 astronauts to the moon. It highlights the skills of these women and their own personal lives, in conjunction with the broader social issues of that time.

The book is moderately long, but it is very well-written and engaging. I liked seeing the descriptions of these towns that flourished during the wartime years and the space race as bustling with life and energy, because with the trends of deindustrialization starting from a few decades ago, I haven't really been able to see descriptions such towns as much beyond shells of their former selves. This also ties in with the discussions of the military-industrial complex and how the formation of these towns during the wartime was a symptom of that phenomenon, which in turn meant that as wartime research facilities and organizations were often temporary, even if they hired black people, allowing them to economically advance to the middle class, those economic advancements became tenuous due to the temporary nature of such jobs, such that when the goal was met (whether it was winning the war or landing on the moon), those facilities would be closed and the employees there would be displaced with few, if any, alternatives available to them. Of course, pervading the book were descriptions of the explicit and implicit forms of institutionalized racism and sexism, whether at work in the form of barriers to career advancement or collegiality/free exchange of ideas, or in the context of daily life with respect to the civil rights movement, sit-ins, et cetera. Not only were those issues discussed in a broader context, but their impact on the specific protagonists of the book was detailed, showing how these women had to deal with so many struggles just to stay afloat while still trying to achieve the same goals to which any other family of any ethnicity would strive, namely, caring for spouses and children, putting food on the table, balancing work and family, and being able to raise children in a safe environment and educate them well; it really helped that the author so masterfully portrayed the mundanity of daily daily life for these women to show how stupid obstacles, like legalized segregation and institutionalized barriers to career advancement, could get in the way of the passion that these women had for STEM. It was also interesting to see that black communities like those in this book were acutely aware of how much more advanced the USSR and other communist countries were in terms of race and gender relations, and actively called out the US on its own failings in that regard (in the context of the US trying to ally with African and Asian countries that used to be European colonies), while these black women, despite being in the middle of such institutional bigotry, kept their heads held high and persevered in pursuit of their goals to contribute to STEM R&D. Related to that, it was also chilling to see how de facto segregation has persisted in education in many places through the US resulting in school facilities that in many poor places are no better than they were several decades ago, and also to see how many of the arguments for white parents sending their kids to private schools at that time were more explicitly about preserving racial segregation in education. Overall, I enjoyed this book thoroughly and would strongly recommend it to people for a clear and engaging account of how NASA and the social issues of the middle of the 20th century became intertwined.

2017-08-14

Book Review: "Genius at Play" by Siobhan Roberts

I've recently been able to read the book Genius at Play by Siobhan Roberts. It is a biography of John Conway from his high school days onward, covering a lot of his work on group theory, symmetries, number theory, and other fields; it does discuss the Game of Life but goes deeper into drawing out the evolution of his response to being solely associated with it, from joy to despair to resigned ambivalence. It also goes through various episodes in his personal life, and frequently switches between narrating recollections of past events and narrating the current events surrounding those recollections themselves. It shows what a whimsical, joyful, carefree, and gregarious man he has been when it comes to math, but while much of the middle section of the book makes this seem like his core personality, the beginning parts about him reinventing his personality after high school, the middle parts about others' (particularly Stephen Wolfram's) characterizations of him, and the end parts about his own feelings about it now (as an older man) make clear that the carefree part of his nature has been more of a facade to cope with his own ego & attendant insecurities.

I really enjoyed this book's progression through his life. It allowed me to once again experience the joys of seeing seemingly inconsequential math ideas that are easy to introduce but hard to truly explain to broad lay audiences, like I did in middle school with my fascination with numbers like pi and the golden ratio and the really random places they pop up. Plus, that juvenile whimsy combined with the human interest in reading about John Conway as a person, as well as my more mature recent interest in deeper philosophical questions, like the physical reality of mathematical theorems (whether they are invented by humans or discovered from the course of nature), or the nature of free will at the human versus subatomic quantum scales. The writing of this book really seemed reflective of John Conway's whimsical and sometimes scattered personality as well as of the people and environment surrounding him, painting a vivid picture of the man in context; I appreciated this to the same extent that I did Robert Kanigel doing the same for Ramanujan in The Man Who Knew Infinity, though in the latter case, Ramanujan seemed to be a rather shy person for whom information had to be gleaned from other sources (also because Kanigel wrote that biography many decades after Ramanujan's death), so Kanigel successfully painted the picture of the people and environment in India and the UK that shaped Ramanujan's personality and the course of his life. Overall, this was a really enjoyable book that went by quickly despite its length. Perhaps people who are slightly more acquainted with the basics of higher-level mathematics may enjoy it more than laypeople (and I would probably enjoy it more if I knew more math); nevertheless, I think the storytelling is quite engaging and accessible to a broad audience. Follow the jump to see a couple other brief thoughts.