Sunday, 4 March 2018

Resynthesizing audio from spectrograms


Sorry about the formatting; it's pasted from a libreoffice document.

Resynthesizing audio from spectrograms


Martin Guy <martinwguy@gmail.com>
Work: July-August 2016; Docs: February 2018.

ABSTRACT

It can occur that the only available source of a piece of music is a JPEG image of its spectrogram. An algorithm is presented to convert such a graphic back into a best-effort approximation of the audio from which it was created.

Here is an example of a source graphic from the case that provoked this work: spectrograms of unpublished samples of electronic music by pioneer Delia Derbyshire in James Percival's 2013 dissertation for his master's degree, Delia Derbyshire’s Creative Process:

Fig. II.4 from Delia Derbyshire’s Creative Process:
“Spectrographic analysis of CDD/1/7/37 (2’49”-3’00” visible)”
CDD/1/7/37 is Singing Waters: “It is raining women’s voices”,
a musical arrangement of Apollinaire’s graphic poem Il Pleut.

This represents 11 seconds of sound in 618 pixel columns (56 columns per second) from 5Hz to 1062.8Hz in 252 pixel rows (so with frequency bins spaced by 4.2 Hz)

It has linear time and frequency axes and is composed of a square grid of coloured points and frequency and colour scales on the left that show what frequencies each row represents and what sound energies are represented by a range of colours.

Algorithm
In brief, we turn the colour values back into estimated amplitudes, then reverse FFT those to create an audio fragment from each pixel column. We then mix these to produce the audio output.

Colour-to-amplitude conversion
We make an associative array mapping the colour values present in the scale to their decibel equivalents by sampling a vertical strip of the colour scale, knowing on which pixel rows it starts and ends and by reading off the minimum and maximum decibel values on the scale. Using this, we map the colour values present in the spectrogram (or their “closest” equivalents on the scale) to create an array representing the energies at each frequency shown in the spectrogram, for each of the moments represented by its pixel columns.

Interpolation between frequency-domain frames
One can optionally reduce the choppiness of low frame-rate spectrograms by interpolating between FFT frames before doing the transform, thereby effectively increasing the frame rate.

Phase
Each reverse FFT, as well as an array of amplitudes, also needs a phase component for each frequency bin, which needs to be chosen to ensure that the sine wave output due to each bin of one frame is in phase with the output from the same bin for the all the other frames.
We do this by setting the phase for a bin centred at f Hz at time t seconds to

random_offset[f] + t × f × 2 pi radians

The constant random phase offset, different for each bin, avoids artifacts caused by many partials coinciding in phase periodically and producing harsh cos-like or sin-like peaks:

                /\                 ,

               /  \               /|

          /\  /    \  /\  or  /| / |  /|

            \/      \/         |/  | / |/

                                   |/

                                   ’

Mixing successive frames

To avoid discontinuities when the output audio changes from the results of one reverse FFT to those of the next, the size of the FFT is twice the number of samples represented by a pixel column, and we then overlap the putative audio output fragments by half a window and fade between them sample by sample to create the final audio data.
The fading function is a Hann window which, being cos squared, has the useful properties that it crosses 0.5 at 1/4 and 3/4 of its width, that each half has 180° rotational symmetry so that the sum of two adjacent windows’ contribution factors is always 1.0, and its endpoints are both at 0. Its bell shape also means that the sound output for the middle half of each window depends mostly on the data from the corresponding pixel column.

In our implementation we centre each fragment of output audio on the time represented by the centre of its corresponding pixel column and mix using a double-width window, so a quarter of the first window extends before the start of the piece’s started start time and a quarter of the last window extends beyond its end, making the total length of our audio output the stated length plus the time for one pixel column.

Results
A program to perform this transformation, specialized for the example graphics, is available under http://github.com/martinwguy/delia-derbyshire in the “anal” folder, file “run.c” with a driver script “run.sh”. The sample input files can be extracted from the thesis, available under https://wikidelia.net/wiki/Delia_Derbyshire%27s_Creative_Process and the audio output from the example spectrogram cited in the text can be heard at https://wikidelia.net/wiki/Singing_Waters

The other spectrograms present in the thesis give similar results, but Singing Waters is the prettiest of them.

Sunday, 21 January 2018

"Do you believe in God?" they sometimes ask me

"Do you believe in God?" they sometimes ask me and my stock answer is that it depends what you mean by "God."

Is that the cartoon God with the long beard and the white robes and the golden throne in the clouds with the angels and the harps and the trumpets? Well, at least I've seen pictures of that one!

Or is that the old-testament Jehovah, sending his favourite tribes rampaging round the desert destroying every village and city they came across? I've read about that one.

Or the one most people spend most of their lives praying to, money, which promises power over other people? That one's for frightened people and, like most promises of power, it leaves you in slavery. I don't believe in that one either.

I once spoke about this to a friend of mine, a friend who spent his spare time doing naive oil paintings and having old bibles rebound in new leather. We were sitting on the kerb of a side street and he confided to me: "You know, Martin, God is an old man with a long beard who walks the streets and when you meet him, if you do what he tells you, he gives you eternal life.

Well, I hope he meets his God, and when he does, I hope he does as he's told!

Saturday, 19 August 2017

Tizen studio is a bunch of broken shit

"Write cellphone apps" they said to me, so I thought I'd have a go. Of the offerings, Tizen, the upcoming cellphone OS, seemed the most promising, promising to compile for X, Windows, Android, IoS and itself. It took me two days to manage to download the Tizen Studio IDE and it's been three days that I've been trying to install the "Native IDE" and "emulator" packages on top of that. After five days fiddling with crap internet and proxies I'm burnt out. Crapware. Slow. Burns 400% of the CPU because written in Java. Doesn't cache partial downloads, so you need a constant fast internet connection to be able to work. I don't have that. It's the usual story: as soon as you touch closed-source software, time stands still and it feels like wading chest-deep in mud. Bastards. Open source it for fuck's sake - we might be able to fix your broken crap. Do I *really* have to figure out how to use it from the command line? After all, that's all I've been able to install! Go figure!

Friday, 12 February 2016

HIdden sounds in Pink Floyd's "Wish You Were Here"

This is a log-frequency-axis spectrogram of the start of Pink Floyd's "Wish You Were Here". The vertical axis covers 9 octaves from 27.5Hz to 14080Hz.

Those bunches of parallel lines are the first solo guitar notes. But what's that diagonal line at the top, a sine tone that sweeps between 2500 and 3500Hz for the duration of the piece until it breaks loose at the end? It should be audible among the higher harmonics of the guitar - maybe try listening to the piece on headphones?

Wednesday, 9 December 2015

Scientific Debugging and "Where should issues live?"

After reading Andreas Zeller's book "Why Programs Fail", I've ben using scientific debugging method while tracing defects in programs.

As I understand it, if you can't fix a bug intuitively in 10 minutes, apply scientific debugging:
- write everything down, keep all generated test files (or, more to the point, the exact commands used to create them)
- follow a cycle of making observations, formulating a hypothesis, devising an experiment that should either confirm ot reject the hypothesis. If it's rejected, you have to formulate a new hypothesis. If it's confirmed it either leads to a diagnosis and fix or to having to refine the hypothesis for a new cycle of tests.

My first attempt was while tracing the causes of timing errors in a buggy FPU, documenting the (many!) steps of the above cycle until I found the cause of the bad floating point results making a test suite fail: that if the result of one FPU instruction happened to be used by a following instruction exactly N clock cycles later and with the right FPU-result-pipeline delays in the intervening instructions, then the result was picked up as garbage.
I'm pretty sure I'd never have got as far as I did without SD.

What I'd like is to integrate the scientific debugging method into the current bug trackers. Where the Scientific Debugging Method is appropriate (variances from desired behaviour with unclear causes), should we look at integrating the finer-grain documentation of the bug's resolution into the issue tracker itself so that, instead of having issues, each with a linear sequence of comments, each would have any number of hypotheses, experiments, observations, with a confirmed hypothesis leading to one or more refined versions of it, and a rejected one having no children.

At present, I'm trying this just by using plaintext and images. Here are two examples with pretty sonic diagrams for sndfile-spectrogram on GitHub: Strange florets in sndfile-spectrogram output and Horizontal comb of shadows in what should be smooth spectrogram output

A few months ago, in a private two-person project without a public tracker, we created a folder in the source tree called "issues", containing text files named "1", "2", "3" and so on, each of which is added to the git tree as an issue-creation commit, together with any associated test files. These text files are then modified as issue investigation and resolution proceeds, and the issue file is removed from the source tree with the commit that resolves the issue. It lacks some things, like being able to see, for example, all resolved issues, though I guess a shell script should do it.

But, hang on a minute, a source code tree's issues *should* be part of the source tree.
If you take a copy of a source tree, you should get the list of known bugs in it too.
Web-based bug trackers like Github divorce the code from its issue tracker and keep the list of issues on their servers. I mean, I think web-based issue trackers are great! But can we think about doing something similar, that travels along with each copy of the code it speaks of?

And can that the operations on that in-source-tree mechanism embody scientific debugging steps in its logic, dealing in Observations, Hypotheses, Predictions, Experiments, Rejections or Confirmations and Diagnoses-Fixes? Is the SciDebug workflow diagram comprehensive enough to make it a mandatory procedure in a bug tracker? If so, it would give speed and power to our collective debugging by forcing everyone working on the project into a "scientific" way of reasoning in their work, as well as documenting the reasoning that led them to their conclusions.