back
83 comments
The baseless log here is just a torsor [0]!

Lots of things are torsors: position, currency values, calendar dates etc. the vales themselves are arbitrary, and translating/scaling them by some value doesn't make a functional difference. Torsors let us talk about these things without needing to make such an arbitrary choice a priori.

In the case of baseless logs, the underlying set is "information units", i.e. log 2 is bits, log e is nats, log 10 is digits, etc. The conversion factors give us the torsor's group, and picking a privileged unit is just a trivialization of the torsor.

The vector division notation is, similarly, encoding a g-torsor in precisely the same way as length units are.

The examples so far are all torsors with abelian groups, but specifying position both requires choosing an origin and a length unit. The group of this torsor is a suitable semidirect product between translation and scaling, which gives a non-abelian group.

Most of the time we just implicitly choose a trivialization, which often causes confusion because it identifies objects with operations on them, e.g. conflating vectors as positions with vectors as translations. The author's treatise on problems with geometric algebra [1] even brings up this point!

[0]:https://math.ucr.edu/home/baez/torsors.html

[1]:https://alexkritchevsky.com/2024/02/28/geometric-algebra.htm...

Using the term "torsor" for that mathematical concept has been a very bad choice, both because the concept does not have any obvious relationship with the meaning of the word and because the word "torsor" had already been used for a very long time in classical mechanics for a very different concept, i.e. for the quantity that must be null for a rigid body to stay in equilibrium (i.e. the pair of a resultant force and a resultant torque).

Unfortunately, in mathematics there already is a long tradition of reusing common words to designate concepts that have no relationship whatsoever with the original meanings of those words. This obfuscates the content of many mathematical books or research papers, because even when they state trivial facts the statements are opaque for those unfamiliar with the specific jargon used in that niche branch of mathematics.

Words happen more than they are chosen, cf. "computer". The term "torsor" in this sense likely comes from the French "torseur" [0], which was used to describe rigid-body motions via a fundamental screw-like action.

The hypothesis seems to be that the idea of affine spaces came out of that theory, for whatever reason, which was subsequently generalized to principle bundles and finally into what we have now. The point is that, at every step along the way, we want to connect the incrementally new ideas to existing ones, and creating a hard break with new, idiosyncratic terminology is itself obfuscatory.

My beef is more with use of the heavily-overloaded words "regular" and "normal" in math, which just seems like lazy naming:

> In the normal extension K/Q, every normal subgroup of the regular representation acts on a normal scheme that is regular in codimension one, whose normal bundle — orthonormal to the regular surface at each regular value — carries a normal operator whose spectrum follows a normal distribution over a space that is at once regular and normal, all indexed by a regular cardinal.

That's like 8 different meanings of normal and 6 different meanings of regular. lol

[0]:https://fr.wikipedia.org/wiki/Torseur

Yeah, see this thread --- I assume these guys haven't heard of the other meaning neither

https://golem.ph.utexas.edu/category/2013/06/torsors_and_enr...

Consider in particular that use of ‘distance’

>I think you can look at adjoint profunctors from the unit category and show that they consist of giving a consistent ‘distance’ to every object, which in a torsor will be represented.

I do know about torsors actually but I didn't think to link it from there. I guess I don't find the term very useful; it feels like things are still hard to think about even after you know it's a torsor!---but also, I think I need to get more familiar with the concept, because the other commenter on here who described my basis-logarithm as a "GL(V)-torsor" really said it much more succinctly than what I was hacking out manually.

Regardless of the terminology, I thought it was interesting because I have never seen the logarithm thought about in that way.

Thanks for the article. I do think your more elementary approach is good pedagogy since the subject is so broadly familiar already. I just like torsors, since they elegantly encode the "arbitrary choice" needed to deal with lots of objects.

Thanks for the writeup!

Thanks for sharing, very interesting. I wonder how this maps to swe
Logs are awesome. I started a math textbook from the 1920's a while ago, and all the calculations relied on tabulated logs, where you would convert the number to a log in a table to reduce the operation's degree, then convert back to the ordinary representation. This would reduce operations like finding cubed roots to division, would could be converted to log-log to be further reduced to subtraction before you would restore to ordinary notation. It feels like you're using a magic wormhole or something when you're doing this stuff by hand, it's really neat.
The physical version of that magic wormhole is called a slide rule.
Another neat application, if a bit simplistic, are these mechanical paper computer that let you figure out your body-mass-index. They are basically two disks with logarithmic scales on them that you rotate relative to each other. Like a slide-rule, but circular. I think you can find them under the name 'BMI wheel'.
Yep, we used manual math + some log tables for calculations in our school exams as late as last decade. Since calculators were not allowed. The exam would be such that you would need the log tables once or twice over the course of the exam. Example: dividing = lookup(a)-lookup(b) and then lookup that in the inverse log (i.e exp) tables.
Got a PDF? I love old books like this.
care to share the name of the said book?
Trigonometry for Navigating Officers by WP Winter

https://www.google.com/books/edition/Trigonometry_for_Naviga...

I found this book because I was a little rusty on my trig and most celestial navigation texts will just throw the PZX equation (and others) at you without breaking down what's actually being done with it on a mathematical level...it's just kind of treated like a magical black box without any discussion, and I'd rather have a complete understanding of what I'm doing and why. Having an application-specific approach also makes it a lot easier to learn.

I'm using it with Norie's Nautical Tables, which has the log tables and a whole lot else:

https://bluewaterweb.com/product/nories-nautical-tables-2025...

I'm sure there are plenty of free PDF's of log tables you can find though.

(I believe they used log tables on boats primarily because it's easier to use than a slide rule when everything is constantly rocking back and forth.)

That's a lot of ways to think about logarithms.

Logarithms are laughably simple once you've fully internalized the meaning of the log function; it simply answers the question:

"To what power must I raise the base to get the argument?"

This is why the output tapers out as you increase the argument; because even if you increase the argument exponentially, you only need a fixed increment in the power to reach that number... So if you increase the argument only by a fixed amount (linearly) instead of exponentially, then it makes sense that the output will grow sub-linearly.

I remember when I was doing algebra with logs many years ago at school, I was applying rules to remove the log from one side of the equation.

Then when I got to uni, I had to revise the rules but it was kind of silly of me because those rules can be trivially derived if you just think about what the log function means. Turns out I had been solving equations with logs throughout school without understanding what they even meant... It's only at university that I actually bothered to learn them.

Actually TBH. I didn't even fully understand powers for some time even though I was doing calculus with them at school. I only fully understood powers once I properly internalized the concept of k-ary trees as a proxy.

It's one thing to be able to apply something, another to understand it. And I think to innovate with something, as a tool, it's not enough to be able to apply it. You must understand it.

A better way to understand logarithms is to start with the original motivation from Napier himself (https://sites.pitt.edu/~super1/lecture/lec44911/005.htm);

Seeing there is nothing (right well-beloved Students of the Mathematics) that is so troublesome to mathematical practice, nor that doth more molest and hinder calculators, than the multiplications, divisions, square and cubical extractions of great numbers, which besides the tedious expense of time are for the most part subject to many slippery errors, I began therefore to consider in my mind by what certain and ready art I might remove those hindrances. And having thought upon many things to this purpose, I found at length some excellent brief rules to be treated of (perhaps) hereafter. But amongst all, none more profitable than this which together with the hard and tedious multiplications, divisions, and extractions of roots, doth also cast away from the work itself even the very numbers themselves that are to be multiplied, divided and resolved into roots, and putteth other numbers in their place which perform as much as they can do, only by addition and subtraction, division by two or division by three.

This is what provides the intuition viz; convert multiplication/division/etc. of large numbers into addition/subtraction of two other smaller numbers. Logarithms as inverse of Exponentiation came much later. Starting with this generally confuses the student since they do not understand the point of it all.

From https://en.wikipedia.org/wiki/History_of_logarithms;

Napier conceived the logarithm as the relationship between two particles moving along a line, one at constant speed and the other at a speed proportional to its distance from a fixed endpoint.

Since the speed is directly proportional to its remaining distance from the fixed endpoint, it therefore is a deceleration, which results in the characteristic "flattening" of the curve.

Further details for understanding the above can be found at Priority, Parallel Discovery, and Pre-eminence: Napier, Burgi and the Early History of the Logarithm Relation (pdf) - http://www.numdam.org/item/RHM_2012__18_2_223_0.pdf

I find it surprising that logarithms and e (a.k.a. Napier’s constant), were developed and discovered only relatively recently in the history of mathematics despite how natural and fundamental they are.

The idea of exponential growth and the practice of charging interest in finance are both ancient. Surely an ancient mathematician would have investigated these in depth and discovered what Napier, Bernoulli and others found?

A while ago I was solving several infinite series of exponentials in the context of a problem concerning the half-life of medicines and I made frequent use of logarithms. That is when I started to wonder about their history.

What made you want to understand it or did it happen upon you in college
This essay needs a type system. Every time it says “log” it should say: log of what, into what?

It’s like audio where people say "dB" as if it answers the next question. Relative to what, measured how, and weighted for whom?

Author should brush up on https://en.wikipedia.org/wiki/Lie_theory

The important properties of the logarithm are structural: we usually do not care about units or bases, except when carrying out an actual numerical computation.

As developed in the article, informally, but somewhat sufficiently, the change of base formula shows that the choice of base is largely irrelevant: different bases give equivalent logarithms up to a constant factor.

The Taylor expansion of exp gives a more intrinsic and general definition of the exponential function. This allows exp to be generalised structurally to many algebraic settings, provided the relevant convergence conditions are met: for example, the complex exponential and its many possible logs, the matrix exponential, and so on…

The first section details how the author thinks of "log N" with no base as an abstract object rather than a number. Or what are you referring to?
I still don't understand why audio dB are negative. That's relative to what? What happens at 0dB?
I think what's going on with the complex logarithm is basically the same as the logarithm that outputs the set of all possible bases for a vector space. The complex logarithm produces a Z-torsor, and the basis logarithm produces a GL(V)-torsor. There's probably some way to represent a choice of branch cut as a part of the choice of the base of the complex logarithm, and similarly the choice of a specific basis as part of the choice of base of the vector space base logarithm.
Interesting, it did not occur to me of those as two instances of the same phenomenon. Although I still find the complex analytic one hard to think about.
Charles Petzold's The Lost Art of Logarithms is a great read (still a work in progress).

https://www.lostartoflogarithms.com/

This looks great; thanks for the pointer.

Charles Petzold's writings are always very clear and in-depth.

>You might ask: if we have a baseless logarithm log(N), do we also have a “baseless exponential”?

Sure we can, with some naive algebra. If we can take log(x,base) and drop the base, then we can also take pow(base,x) and drop the base. Since bits=log(2), then pow(bits)=2. You can probably connect it to the reverse of things, like integrals.

Also, for fun, I'll play with some notation tricks.

  log(freq) = pitch
  freq = pow(pitch)
  octave = log(2)

  400*Hz = 100*Hz*4  // the frequency 400 Hz equals 4 times 100 Hz
  log(400*Hz) = log(100*Hz) + log(4)
  log(400*Hz) = log(100*Hz) + 2*log(2)
  log(400*Hz) = log(100*Hz) + 2*octave
  log(400*Hz) = log(100*Hz) + 2*octave  // the pitch of 400 Hz equals 2 octaves above the pitch of 100 Hz

  cent = log(2)/1200
  A4 = log(440*Hz)
  B4 = A4 + 200*cent  // the pitch B4 equals 200 cents above A4
  B4 = log(440*Hz) + 200*log(2)/1200
  B4 = log(440*Hz) + log(2^(2/12))
  B4 = log(440*Hz * 2^(2/12))
  pow(B4) = 493.883 Hz  // the frequency of B4 equals 493.883 Hz
I like the intuition that baseless logarithm notation gives, and it also avoids needing to choose a specific reference point. I can also directly calculate by choosing an arbitrary base:

  pow(log(440*Hz) + 200*log(2)/1200)
  exp(ln(440) + 200*ln(2)/1200)
Hah, I can use this to give decibels an actual unit.

  dB_P = log(10)/10
  dB_F = log(10)/20
  log(10*V) = log(V) + 20*dB_F  // the level of 10 V equals 20 dB more than the power level of 1 V.

  SPL = 20*10^-6 * Pa
  hearing_damage = log(SPL) + 90*dB_F  // hearing damage occurs over 90 dB_F above SPL (neglecting A-weighting)
  pow(hearing_damage) = pow(log(SPL) + 90*dB_F))
  pow(hearing_damage) = pow(log(SPL) + 90*log(10)/20))
  pow(hearing_damage) = SPL*pow(90*log(10)/20))
  pow(hearing_damage) = SPL*31622.7766  // the pressure of hearing damage occurs above 31622 times SPL
  pow(hearing_damage) = 0.632455532 Pa  // the pressure of hearing damage occurs above 0.632 Pa
Very helpful!! Imagine combining the goofy list of decibel suffixes into a uniform notation. Write the logarithm first so the + or - stays in the same spot.

  log(reference_unit) + value*dB_F (or dB_P)
  log(reference_unit) - value*dB_F (or dB_P)
https://en.wikipedia.org/wiki/Decibel#List_of_suffixes
True, I guess you can just 'curry' exponentiation and say that's a baseless power. I couldn't find a clean notation for it so I gave up..
The same idea comes up in physics. In quantum physics, the action S appears as the logarithm-like quantity behind the amplitude e^iS/(h^bar). In statistical mechanics, entropy is the logarithm of the number of possible microstates Omega : S = log(Omega). Although the concepts come from different parts of physics, they both reflect the same principle: using a log as a way to turn multiplicative relationships into additive ones.
I can't believe he called normal logarithms 'based'
The term "baseless logarithm" is really nonsensical and using it would be a great mistake.

Nonetheless, where the author of TFA is correct is that logarithms are a single physical quantity, like length, area or volume, and that choosing the so called "base" is choosing the unit of measurement for logarithms.

Logarithms are included in the dimensional formulae of many derived physical quantities, e.g. for describing the attenuation or amplification of waves during their propagation, where one uses quantities like logarithm per length and logarithm per time.

Changing the "base" of logarithms modifies the numeric values of all derived physical quantities exactly in the same manner as changing any other fundamental unit of measurement, like the unit of length or the unit of time.

Like for any physical quantity, the complete value of a logarithm is independent of the unit of measurement, because it is the product between the numeric value and the unit of measurement. When the unit of measurement is changed, both the numeric value and the unit are changed and the product stays the same (i.e. the logarithm corresponds to the same ratio, regardless what base is used to compute a numeric value for the logarithm).

Nowadays, the unit of logarithms is normally chosen between the octave (binary logarithms), neper (hyperbolic logarithms) or bel (decimal logarithms).

The units of measurement for logarithms are not the bases, but the logarithms of the bases, which is why e.g. the value of the number "e", the base of the hyperbolic logarithms, is never needed in any computation. The only values that are needed are "ln 2" or its inverse "log2 e", which are used to convert the numeric values of logarithms when the unit of measurement is changed between those corresponding to binary logarithms and to hyperbolic logarithms (a.k.a. natural logarithms, but there is nothing more "natural" about hyperbolic logarithms than about any other kind of logarithms).

"Baseless logarithm" is not nonsencial. Given that:

    d(logₐx)/dx = 1/(x log(a))
a baseless logarithm is simply a family of functions with similar properties. Perhaps it might be clearer if the author said something like the "logarithm property" rather than "baseless logarithm" but that's nit-picking and debatable.

As for changing the base changes the numbers, I have to wonder if you've done any advanced linear algebra or, more specifically, tensors. The whole point of a tensor is that it operates the same on an object regardless of the basis. Put another way, if a and b are two representations of the same object with different bases then T(a) and T(b) are equivalent if T(x) is a tensor.

My point is that any numbers are an arbitrary choice and they don't define the underlying structure. The author here is talking about logarithmic structure.

This btw is why you learn about different bases in linear algebra and converting between them. Or even polar coordinates vs cartesian coordinates (in high school, for some reason). They're priming you to learn about structure. You get to groups and learn that group A and B are isomorphic they have the same mathetmatical structure.

Even when the numbers change.

All this would be way more interesting if it actually helped to demonstrate a novel mathematical fact. Right now it's more like notational play.
I read this kind of essay as a certain part of the arc by which new thoughts are formed: an act of large-scale pattern matching, laying out a bunch of cases which resemble each other, searching for the essential basis of the resemblance.

To post such a pattern allows the thought process to become distributed. Perhaps someone else will see the insight.

I happen to think that novel facts and theorems and proofs are way overrated. If you find a new fact it just goes into the giant pile of facts that are sitting around uselessly. The useful progress in math is comes from "refactoring" efforts to make things simpler and more intuitive.

I don't mean that this is necessarily the case, but that it is where we are now: we have found ourself in a situation where we have way too many facts and not enough simple perspectives that make them useful and accessible.

Just my opinion, though.

IIRC, Knuth use lg for logarithm base 2.
Does this answer the question of why we see hyperoperations until exponentiation in physics, but not higher?
I think that's more about integrations/differentials not producing them (generally speaking). Physics likes to deal with integrals and differentiation as you calculate change over time or over spatial dimensions.

Eg. the integral of x^10 is x^11 / 11 + c. No hyper-operation appears and it's just another exponential (with a division).

The integral of log(x) is xlog(x) - x + c. So still basically just a logarithm

Even the integral of 2^x is just 2^x / log(2). Still basically the same thing.

There's no easy way to pull a hyper-operation out.

Wasn't there some scientific paper recently that proved that every operation can be represented as a logarithm? Like, the same as every logic gate can be derived from NAND gates
Was it this exp-minus-log arxiv paper?: https://arxiv.org/html/2603.21852v2
This sentiment (not necessarily the content) is what I'm striving to communicate with Mag World[0] (website and podcast so far).

[0] magworld.pw

tl;dr: Being a homomorphism from a multiplicative structure into an additive structure isn't enough to grant it the logarithm title.

Although logarithms are certainly ubiquitous in mathematics, I don't think that the mappings that the article's author identifies as logarithms are appropriately viewed as such.

I can't endorse viewing dimension as a logarithm. It appears superficially logarithm-like because we typically (and somewhat unfortunately) write the direct sum of n copies of a vector space V as V^n rather than nV. Writing nV, we simply get the dimension identity dim(nV) = n dim(V). Writing nV instead of V^n also conveniently frees up V^n for the tensor product of n copies of V, with corresponding dimension identity dim(V^n) = dim(V)^n. So I don't think there's any "multiplicative-to-additive" business going on here at all.

Also, I don't think it's advisable to view the p-adic valuation ord_p as a logarithm, even though it's a homomorphisms from the multiplicative group of the rational or p-adic field into the additive group of the rational field. In fact, in many number theoretic contexts, the ratio log_p/ord_p is of particular interest.

I think a good rule of thumb for viewing a mapping as some kind of logarithm is that it has to have some relation with the Taylor expansion of log(1 + x) around x=0. Being a homomorphism from a multiplicative structure into an additive structure isn't enough to get the logarithm title.

Look, the whole thing actually makes sense and the core idea is pretty cool because it's true that a lot of stuff in math looks identical. But in my opinion this is way too much of a macro-level overgeneralization and you risk throwing everything into the same pot, which ends up diluting the actual point of things.I mean, if you take a hammer and a meat mallet, at the end of the day they're both chunks of metal used to hit stuff, but if you bunch them together without making any distinction, you lose track of why you use one to drive nails into a wall and the other to prep cutlets.Saying everything is just one big logarithm is a nice mental exercise, but I feel like it flattens out the differences too much and makes you lose the practical utility of the individual math tools, which are meant to solve completely different problems.
I'm a programmer so to me this brings to mind the idea of classes and subclasses. A program is implemented by having a set of classes. The classes can be organized into a class-hierarchy where they inherit methods from their ancestor-classes.

Now assume originally you did not have the feature of inheritance in your programming language so you would just create all the classes you need without orgnizing them into an inheritance-tree. Then you upgraded to a language that doe shave inheritance and you wanted to refactor your program to omit duplicate definitions of methods.

What kind of class-hierarchy would you come up with? There is no single way to do it. Some ways are better than others. There migh be more than one optimal way.

Same goes with generalization general, it is part of the language we create to describe things and there are many different languages we may come up with, some simpler, some more difficult to understand.