No, we hear things as if they were a combination of sinusoidal waves (mostly). That's only one way to represent them; you can just as easily represent them as a series of impulses. Mathematically they're all equivalent models, and none of the models accurately represent acoustic transfer.
You can say we perceive things as if they were a linear combination of sinusoidal waves because we have a series of hair cells in the cochlea. Each bundle of hairs responds to waves within a very narrow frequency range, so that's what we mainly respond to. All of our audio perception is filtered through this process first. Since the hair bundles have a specific audio response, they don't respond to instantaneous frequencies- there has to be a wave of a certain duration for us to perceive a frequency. That's even independent of the mathematical fact that shorter wave pulses are more indistinguishable from white noise. Also, all of this is ignoring the other equipment of the ear, like the eardrum and the bone lever[1] that transmits sound to the inner ear.
We have a lot of brain circuitry that picks out specific features, which means we can recognize things like square waves or impulses. I don't know much about this stuff but it's important for things like cochlear implants. AFAIK it's similar to some optical illusion/visual perception phenomena- things like how we perceive magenta as a single color, even though it's just two colors together, or how we perceive orange as a distinct color from brown even though it isn't[2].
> The more you have it's as if there are more instruments playing on top of each other, in music they're called "harmonics".
Symmetrical clipping and square waves both create odd-order harmonics (3x, 5x, 7x, 9x, etc) which are very unnatural. The brain is good at picking this kind of thing out, and while I find it enjoyable (eg chiptunes), it's not really comparable to normal harmonics in music. It's certainly not something you can simplify down to just the component sinusoids. The perception of sounds has way more to do with stuff happening in the brain than in the air.
[1]: https://en.wikipedia.org/wiki/Ossicles#/media/File:Slide1ghe...
What do you mean by "unnatural"? A clarinet has mostly odd harmonics[1], for example. Clarinets don't occur in nature, I suppose, but they're not using electronics or anything either.
[1] https://newt.phys.unsw.edu.au/jw/clarinetacoustics.html#harm...
I plan to share on the blog at some point https://omnisplore.wordpress.com/category/music-creative/
> No, we hear things as if they were a combination of sinusoidal waves (mostly). That's only one way to represent them
Well, but the reason we do our audio processing this way is that the Fourier transform our ear puts out is a high-quality description of the sound. If it didn't work well, we wouldn't do it.
> or how we perceive orange as a distinct color from brown even though it isn't
Why orange and brown? You can say the same thing about red and blue.
Conversely, if orange and brown weren't distinct colors, it wouldn't be possible for us to perceive them differently, and yet we do, very consistently. What exactly are you trying to say by "a distinct color"?
Anything a human can perceive as sound can be simplified down to just the component sine waves.
> You can say we perceive things as if they were a linear combination of sinusoidal waves because we have a series of hair cells in the cochlea.
No. We say this because any waveform can be reproduced by a sufficient number of sine waves at different frequencies and amplitudes. It's not perception, or the physics of the human ear. It's math.[1]
[1] http://astro.pas.rochester.edu/~aquillen/phy103/Lectures/D_F...
Math does not have anything to say about sums of sine waves being the true form of functions. The wavelet transform is just as physically representative. In fact you can generate any number of Hilbert bases for a given function. They are ALL equivalent.
The Fourier transform has a particular relevance because ears do something similar. Hair cells signal the brain while they are sensing vibrations within their particular frequency range. You could have different ears that signaled when they saw a sharp rise or drop in pressure and worked completely differently, but they would sense square waves just as fine. Or you could have ears that directly sense and measure air pressure. Instead, our ears sense that there are vibrations at 1 Hz, 3 Hz, 5 Hz, 7 Hz etc. and the brain interprets that and realizes it is a square wave.
The concept of a real Fourier transform is just as unphysical as a real square wave. A real Fourier transform requires a perfectly defined sound pressure at every instant; because air and sensors have inertia that is not possible. There will always be lag, so any change in frequency is not perfectly transformed or represented from/as a sum of sine waves.
[1] https://xiph.org/video/vid2.shtml It's 24 min long and, imo, does an excellent job at explaining the role of digital sampling frequencies, bit-depth, and the misleading "stairstep" representation of digital waveforms.
The difference between digital and analogue gear is that the imperfections in analogue gear make them interesting. i.e. instabilities in VCOs, the addition of harmonics etc. I am a big, big fan of old analogue gear and so I am definitely in the camp of 'we're not there yet' with the emulation of analogue gear.
But, the idea that a digital signal can't replicate it, is nonsense, because obviously an analogue signal can be recorded digitally and perfectly played back. This is often the myth that is propagated, that the 'stair steps' somehow mean a loss of fidelity.
Interestingly, and relevant to this discussion, I am yet to find any digital plugin that manages to do saturation properly. Plugins are often, seemingly, just using random number generators, rather than anything more complex (like the AI models you link to); Nothing comes close to my Thermionic Culture Vulture, or any of my other valve based equipment at creating that 'fatness'.
I think that the AI approach is interesting, it does make me wonder how far we're willing to go to reproduce this in software, when the real thing is a handful of wires, resistors, capacitors, and transformers. It seems like trying to implement the ARM instruction set in Javascript. How much processing would a simple saturation plugin need? How many could one computer run?
Something like U-He Diva for example can sound as analog as anything analog that I've heard. Then if you get into wavetable or granular, you can get sounds that are impossible to achieve with analog, which are in my opinion very "rich and interesting".
I suspect that the people who don't agree may have heard a digital synth many years ago and are unaware that the technology has improved markedly since then, which isn't surprising because there's a lot of money in it.
Just my opinion, but modern analog emulation sounds just as good as real analog.
Nothing wrong with enjoying the workflow of analog though, although personally I dislike that workflow.
Perhaps, for relatively simple sounds, or sounds that have little movement. But, the moment there's any filter movement then it's obvious. The filter really is the key, I think, when a digital synth outs itself. Often there's a 'digital edge' to filters on digital synths that don't come close to the 'smooth destruction' that happens in the analogue realm (really, really hard to describe with words!). But, in A/B tests between digital repros of classic synths and the classic synth, that's nearly always where I can hear the difference. But, it's also a major difference, because often the 'sound' of an analogue synth comes from its filter.
If you're looking for that stand out sound for a track, the one everyone goes "I love the track with that noise in", then analogue is where it will come from, the quickest and easiest.
This obsession with replicating the analogue sounds in software is really tedious though. I prefer digital synths that do stuff that the analogue synths can't do, like Omnisphere - nothing touches that, and it's a perfect compliment to the real analogue sounds. Let's keep the analogue realm, doing what it does best, and then let's advance the possibilities within the digital realm, looking forward rather than back.
The point I was trying to make is that I feel that digital is definitely good enough to do professional-sounding "analog" performances without people going, "definitely sounds digital" - unless digital is the sound you're trying to achieve.
>If you're looking for that stand out sound for a track, the one everyone goes "I love the track with that noise in", then analogue is where it will come from, the quickest and easiest.
I disagree, but I also recognize that nobody will "win" this discussion, because it's a matter of taste - like asking for a consensus on which tastes better, chicken or fish.
You enjoy your analog synths, and I'll enjoy badly playing my digital synths. I'm more of a modern wavetable fan than analog sound anyway.
We're both music lovers - nobody's forcing anyone to use anything.
Statements like this need to be prefaced with qualifying disclaimers, like "IMHO". You have to accept that your opinion is completely subjective. I've seen a variety of blind tests over the years in which listeners were unable to differentiate between analogue hardware and digital simulations. I've become increasingly impressed with the quality of virtual synthesisers over the years. At this stage, while I do still often prefer to use my analogue synths over the VSTs I own, that's mostly because the hardware is simply more fun to experiment with.
You're right in the strict sense about analog random bias being hard to reproduce, but those are largely irrelevant when it comes to the important characteristics of the sound. In an A/B test, it is impossible to distinguish a well designed digital audio simulation from a real analog device.
Could you go into more detail on why you say this? A real life square wave is square, possibly with some ringing or overshoot in the flanks (but not necessarily)
The best you can get is a tight curve into an extremely steep "triangle".
The amplifier driving the speaker has finite slew. The speaker itself is constrained to move at finite speed and cannot accelerate instantly. The air has a finite slew rate - the speed of sound.
Stackexchange has a discussion on trying to determine a maximum frequency in air here: https://physics.stackexchange.com/questions/23418/is-there-a... which leads to all sorts of useful sub-discussions. Attenuation depends on frequency; the higher frequency harmonics are more subject to attenuation in air, as well as reflecting off surfaces and self-interfering.
What you might have is that the air might disperse the wave so it gets distorted but still with the same components https://en.wikipedia.org/wiki/Acoustic_dispersion
Since air pressure does not typically produce a height difference in the atmosphere, standard sounds in air are not subject to dispersion and act very differently from water. There is obviously still a limit to the rise/fall time of a square wave, when the transition is more like a shockwave front, but that's a very square wave.
Yes it's easy to think they behave the same, but they don't.
(below is the original 440Hz square wave; above is the recorded sound)
The recorded waveform looks nothing like the original, and it sounds very different too: the original is much harsher, although the pitch sounds exactly the same, as expected.
Might be my crappy microphone? Or maybe the sound is being filtered somewhere along the way?
My best guess for the large attack showing up there is not effects from the microphone, DAC, amp, or anything but the actual speaker. Good audio measurements are hard to come by thanks to all the snake oil, but square wave measurements are common when there is data. The physical models of transducers aren’t trivial, but in a single broad stroke the answer is “physics” and “spring-mass-damper”.
https://www.innerfidelity.com/images/AKGK701.pdf
The decay is expected, as the pressure around the microphone can only temporarily increase before the pressure wave disperses into the room. As an extreme example, if you turned an entire wall into a speaker and played a square wave, you would be able to make a much more square shaped waveform. There’s no replacement for displacement baby (see: subwoofers). If you want a more square waveform, try using a headphone pushed right up to the mic and going to a higher frequency (like 1 kHz).
Edit: If you want to skip the trip down the rabbit hole: I think the end-all for audio quality are sealed in-ear monitors (IEMs). Low group delay from short distance from transducer to eardrum, and sealed enclosure for good bass response (see: square shaped low frequency square waves).
The point of the experiment is to demonstrate that sound waves in air are not at all like water waves and sonic square waves do, in fact, exist.
But it is true: if there was a plate that could move like a square it would not mean all air in the front and back of it would also move like it.
The air compression would move like a sine instead of a square.