back
115 comments
This is primarily architecturally interesting in my opinion. Output songs have unusual noticeable artifacts, and I would guess they become more noticeable the more you listen.

That said, wow. An end to end FAST architecture that can infer a 4.5 minute song in 10 seconds is a compelling thing. I didn’t see if we got open weights, but my guess is that this is not crazy challenging to train, and some v2/v3 versions of this are likely to be good-to-very-good.

The huge missing issue is direction. Songs are way more than just a 10 second style reference and lyrics. Even the most generic pop song from the 90s had recognizable choruses some repeated bars and some ebb and flow to the song that connected to the lyrics to make it interesting to the human ear. Right now the generated songs, as you noted, somewhat glitchy lyrics over a bland backing track that just sort of goes at one speed and note for the whole of the lyrics.
The style matching is interesting, but there's no song structure. There's no identifiable chorus in any of the demo songs.
I find this very surprising, because it's one of things I'd have expected a diffusion model to have a chance of achieving.

I suppose it might be because it's latent diffusion.

That can probably be a style in itself (if we kept exploring in these directions)
No, it’s not a style, it’s by definition an incomplete song.
If I am to retain any interest as an amateur music writer without proaudio engineering skills and equipment, but with a day job, , I want tools that help me enact MY vision to reality. That means multi tracking, ability to hum or score a melody and have it transfer to musical instrument, ability to enter existing tracks, provide a temporal segment for diffusion, and ask it to 'generate a counterpoint to the melody with strings, etc. The most exciting possibilities of this is enabling talented writers with day jobs, not one click song writing.
^ THIS

As an amateur musician, I'd like tools that help me be more productive musically - those that complement my skills (whatever they may be). All the things you mentioned above, namely, ability to score a melody via a simple hum, transfer to various instruments, generate proper responses to calls, generate melodies within a framework, etc., all these would be super valuable to me.

I'm an OK guitar + bass + keyboard player, I'd LOVE to have an AI assistant that accompanies along. That would make my own jammin' so much richer.

I dont think we have seen the end of AI-driven tools in music-tech yet. I'm cautiously hopeful.

I definitely see this happening. Music generation has lagged behind image generation but is following more or less the same path. Early image generation models were completely unconditional; all you could do was sample an image. Then coarse conditioning methods such as text prompts and depth images came along; then additional tooling to tune images in a more fine-grained way.

That said, there is a difference to images in that music also has a "symbolic" level to it that is closer to text than images [1]. There's other work out there that uses LLM-type tools for direct melody generation (no audio). And of course, there's lyrics. I do expect commercial tools to start integrating all these capabilities gradually, it's just a matter of time.

[1] I guess there's also vector images (like SVG) - I've seen work in generating those as well, though it's less mature than directly generating pixels.

I'd be surprised if the people who are writing these understand what this means.
The request is valid; you just need the right tools for the job.

Story Jam lets you design chord progressions without needing to know about music theory, instead offering intuitive terms like "lightness", "darkness", "drifting" and "roaming". They mean about what you think they mean.

https://storyjam.tenpens.ink

I'm planning a "Show HN" post for tomorrow morning EST with more details. But you can get the sneak peek here :)

Yeah, I'd think that it will take commoditized generation tools that existing or new composition multi-tracking tools could incorporate. i.e. FLStudio plugin
The people writing the one-shot tools are living a pipe dream and are riding the hype wave. One-shot AI music will have a short amount of interest based on its novelty, but the very next generation of humans will revolt against it as a cringe decision of the old guard. Form there it might finally be applied more realistically as an aid to human expression instead of a replacement.
one-shots are essentially toys
Goodness, the music that is produced has almost no discernible time signature. I don't know if my brain is faulty, but I find it extremely annoying to listen to.
It’s AI slop, that’s expected.

“Pop punk with prog rock time signatures“ is a funny idea, but it’s not interesting to listen to when there’s obviously zero intention behind it.

Cool. Obviously needs some work. Lots of artifacts. Something to build on though.

Lots of sour grapes comments from folks. Too bad. Not what I expect out of Hacker News. Glad people are pushing the technological envelope and exploring this space despite the strong negative emotions.

the "prompt" is the 10 seconds original audio file + the lyrics, right?

absolutely crazy

[flagged]
"Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something."

https://news.ycombinator.com/newsguidelines.html

This is why we never should have invented the phonograph. People who want to listen to music can just buy a record, making it literally impossible for them to perform an activity humans ENJOY doing. Without it everyone would surely be making all their own music, and nothing valuable would be lost
The ability to record has led the greatest expansion I musical artistry in human history.

Ty it don’t think peasants were listening t to Bach, do you? Only the extraordinarily wealthy could afford to have music as anything like an every day thing.

There are like 6 core activities that bind humans together: shared creation of food, myth and music; co habitation, protection, child rearing.

We've done these things ourselves for hundreds of thousands of years. As we are increasingly convinced to buy them for convenience we loose the very things that make us know our connectedness.

So ya, there are real problems caused by the convenience of technology

I don't think there's any stopping it, unfortunately. The internet is too good at "optimising" content. The future is Mr Beast, Instagram hotties and 6 pack guys, tiktok morons and onlyfans. Be happy, the market has spoken.
People that never considered the value of artistic process until it was the topic du jour unilaterally decided that it was inefficient, oppressive, complex, frivolous, and unfairly inaccessible to those that hadn't put any sustained effort into developing theirs. If you didn't understand what they don't, you'd realize that companies spending billions of dollars to create tools that make cheap simulacra of artists' work to sell them at a loss to crush them in their own markets was merely the natural progression of artistic praxis. Despite it being economically unsustainable and clearly only cheap until it craters the value of artistic skill, these tools have democratized creativity. Instead of creation only being available to those with the interest and willingness to practice and develop their artistic sense, process, and skill, they're now broadly available to anyone willing to pay money for a subscription service that will obviously soon be a hell of a lot more expensive, or shell out a few thousands dollars for a top-tier video card that you almost certainly already have in your gaming rig, anyway. This is silicon valley progress and if you don't like it, you're a communist.
Related: why are programmers racing to make the perfect AI coding tool? It's an activity many programmers enjoy, and more importantly, if the pace continues, they will likely be automating themselves (or at least a large portion of programmers globally) out of a job.

Granted, many people are benefiting from these tools (myself included) but at some point a lot of us are going to have to find a new job (assuming the progression continues unabated), and I'm not sure what new jobs are going to exist when LLM coders replace many or most of us.

Dynamic music generation for interactive media seems like a good reason?
Not everyone enjoys composing music, and for a large group of people paying an artist is not an option. There's a lot to critizise about current AI tech, saying of all things this has no net benefit seems like the wrong thing to call out, and incredibly short sighted for HN.
You can say the same about creating images, writing text and coding. Yet here we are with all kinds of AI tools to exactly that…
Generate track, extract the pieces I like, build a real track using those pieces with all the other samples I use.
>You are automating an activity humans ENJOY doing.

There's at least an order of magnitude more people who enjoy making music than there are people with the actual skill/talent to make music. Music generation AI is an absolute blessing to the untalented among us who'd love to make a song in a certain style or with certain lyrics but lack the time, talent or ability to do it ourselves.

why make this- because people like music? I want to use it to make my own music, according to you, I can't because it deprives some /real/ musician of making money? what an insane argument- ban singing unless you're in a choir?
Some humans honestly enjoy automating stuff. We wouldn't want to be taking away something that humans enjoy, would we?

I'm a musician myself, but I sadly suspect that most music made today "benefits humanity" very little... Is music making always a net positive? If nothing else, these tools will allow more music will be made.

because copyright? these AI tools are great for adding a music to your YT video w/o worrying about violating copyright.
rap producers are running out of samples to use for new tracks. Everything vintage that can be used has been used, sampling real estate is pretty much dried up. Why would you take away a project that brings a net benefit to humanity?
None of this is music. It is noise that sounds likes music. Pretty analogous to how AI slop is not information, but just words that are arranged to look like information.
Makes me think of the demons from Frieren who are monsters that have learned to speak and act human to hunt people better.
Business hates creatives. They'll do anything to automate us away.
Business doesn’t hate creatives, and is not specifically targeting creatives to automate them away. Any job that can be done as good for a lower price or better for the same price is going to be a target.
whatever can be automated isn't "true" creativity. these models merely generate an average music, but the outputs of creative musicians always stand out.
If I was a business I'd "hate" creatives too, and I'd also want to automate them away. The costs of producing (truly) creative works is utterly bonkers, and so are the risks associated.
It’s just combining sample WAV files without human coordination, talk about a lame-ass achievement. It’s already easy enough to set BPM and load in files in Ableton and warp them into unison, from what I heard this is basically just that with”HOORAY FOR AI” slathered as a veneer on top.

If you think I’m being harsh, I have my reasons as a professional musician to critique these things in an unflattering light because they are my competition. Thankfully actually “generated” AI music is trash. Copyright is problematic in the US, I admit, but tech bros using copyrighted material to train programs to put us out of business - without paying a penny which even Spotify doesn’t per stream - yeah, I’ll have some disdain about this scenario and I feel it’s justified.

Just because you can doesn’t mean you should.

Sorry no. Here on HN, your having a vested interest in some market makes your opinion entirely invalid. That is, enless you're interested in one of the correct markets such as software or AI services.
One thing that strikes me about almost every AI-generated track (from academic or commercial generators), is that even if it's often "competent" - in that it has reasonable melodies, chord progressions, etc - is how average it is. Mediocre, taking the term literally. In a way that also highlights cliches and crutches that are common in human-made music. Somewhat reminiscent of GPT text that drones on and on in a grammatically correct way but conveys little of interest. This is of course not unexpected, given how these models are trained. I wonder if this will have an effect of pushing (human) musicians to be more experimental - to move away from the conventions that are now just a click away for anyone.
Yeah-- in a professional workflow, at best, these tools are for getting ideas rather than creating output that will be used directly. Lots of folks use them for actual creation because they're just so enamored with the ability to create vaguely technically competent output from text, but they're all pretty much a bee-line to mediocre, and overcoming mediocrity is absolutely the most difficult part of working with AI output. The same is true with text, as you mentioned, and image generators. As Charles Eames said, "The details are not the details. They make the design." Well, these tools suck with details, and details convey character, perspective, message, meaning, etc. Surely the tooling will improve this in years to come, but it certainly hasn't yet.