Need help?
<- Back

Comments (79)

  • hypersoar
    I mentioned this in another comment thread, recently, but I value art, including music, for the fact that another human being(s) created it to express something. I have zero interest in it if it's AI-generated. The thought of endless content created from the void for me to gorge on like oreos fills me with nothing but despair.
  • stevage
    It's really depressing that I have gotten into music composition right when this is happening.All of these projects are basically just destroying entire fields of human creativity.It doesn't seem so bad at first when the quality is mediocre, but as soon as they get good, then any human composition will be completely devalued, despite the enormous amount of knowledge, skill and effort that goes into producing one.A composer who writes a piece of music themselves, instead of being recognised and admired basically becomes the equivalent of someone who bakes their own sourdough instead of going to the supermarket: a curiosity, someone who does things the hard way for their own amusement.Graphic artists and bloggers were the first to suffer this, but it feels like it's just going to happen to basically every type of creative expression that can be represented digitally. It's just a question of how many months until that happens.
  • vunderba
    This seems to be one of the few open-weight models that isn't just another "text to audio" generator. You can take existing sheet music, convert it to ABC notation, and then experiment with various covers.I don't have a ton of time right now to really mess around with it, but I took 15 minutes to run it on my Mac M1 with an old melody I wrote back in high school for an RPG I was writing in VB6 (dating myself a bit here).If anyone wants to hear what it sounds like:https://mordenstar.com/other/yue2-tests
  • tom_vidal
    We are creating animated corpses out of our culture.The point of art isn’t to generate beautiful patterns. It’s to create meaning. And you can’t have that without human care and intent.Furthermore, that care needs to be intrinsically interwoven into the patterns that comprise the work. That’s what craft is. The mastery to generate exquisite patterns, and the care to infuse them with meaning at every level.Without care, in a Heideggerian sense, we are left with something which tastes sweet but is completely hollow, perhaps even illusional.
  • dcsommer
    For an informed take (from a professional musician, composer, and arranger) on where this is going for musicians and the music industry, I really recommend Adam Neely's video essay "Suno, AI Music, and the Bad Future": https://www.youtube.com/watch?v=U8dcFhF0Dlk
  • euleriancon
    I've been waiting for symbolic music generators for ages. This is so nice for musicians. Being able to intelligently translate parts, generate additional parts, generate sheet music from audio, generate variations on pieces, making pieces easier or harder.So many possibilities.
  • dr_kiszonka
    The vocals in the few songs I listened to are so very clearly AI. I find Suno's much more credible in many genres.If anyone has tried Suno's Studio 2, I would love to hear your take on it. At $24-30/month, it's approaching Ableton Live 12 Suite rent-to-own territory, without the "own" part.
  • mg1c
    Taste is still what separates AI output from human output regardless of whether it's music or software, IMO. Thinking in modes vs. functional harmony, how you hear a chord in context, which voicing you reach for etc. are things that can hardly be abstracted into strict music theory, let alone pattern-matched.I imagine you can get pretty close to replicating an artist's sound if you specifically train on their output, but at that point it's the same question of whether you're satisfied with abstracting away the things that make writing music/software interesting to us (with music the answer's a lot more obvious since the point of writing music has always just been... to write music).
  • artmanja111
    I say this as someone with no skin in the game, but I don’t think it’s necessary for AI to replace human creativity in music. Instead, it can be a tool for augmentation.There has almost always been backlash against new technologies in music. Ultimately, musicians adopted those technologies and continued producing great works.I recently dug some 20-year-old CDs out from storage that were full of riffs and chord progressions for songs I never got around to completing. I fed them to Suno as a sort of “what might these have sounded like if I’d finished them?” experiment. I found the results quite impressive and plan to publish them some time in the near future.None of the tracks sound like marketable products to me, and I certainly don't think AI is going to generate the next Bohemian Rhapsody -- but it definitely could be a great thing for curing writer's block.
  • philip-b
    That example "Comedy" sounds like the most soulless synthetic example of "music" ever. Like the music analogue of a PR-speak article of a big corpo generated by an LLM. It's still technically impressive but it's 0/10, absolutely empty of soul.
  • voidash
    To add the fact that these AI models, readily help with soul draining music generation finetuning if you will ask it. it will scrape youtube for audio, find ways to defeat provenance and all the other shenanigans but won't help a genuine cause because it wasn't intendedThis is baffling to me. Idk what alignment means. Perhaps brainrot and endless mischaracterization of creativity is one of them
  • Tsarp
    I wish folks spend more time on agentic music genertion instead of a single black box prompt to music.Similar to coding, where you are have a bunch of tools that can do the drudge work for you (the mundane things in mixing mastering for example), but still retain full creative freedom and expression.I am not articulating this well, but basically transitioning from IDE->Coding harnesses what is the DAW-> ??
  • sexy_seedbox
    With this local audio model, I can finally make proper covers in different genres and change lyrics.Previously: https://news.ycombinator.com/item?id=46641627
  • yellowapple
    It's interesting that the “Classical Music” sample ended up with a style prompt of “Contemporary Christian”.It's even more interesting that the “Contemporary Christian” music it generated was specifically Mormon music (specifically about the exile of Lehi and his family from Jerusalem into the wilderness, during which they reach the ocean, build a boat, and sail to the Americas), but in the style of contemporary Christian rock.
  • corysama
  • ipsum2
    The model was trained on CC0, "royalty-free" music, so the outputs are going to sound a bit bland and generic. Model is not currently possible to finetune without the semantic tokenizer.
  • bit-rot
    The majority of people, I think, don't want to listen to AI music or view AI art. Systems like this will produce muzak for ubers, cafes, phone queues, advertisements. Some artists will make hyperreal work by sampling it. But to listen to these "songs" willingly is not something I expect most people will choose to do.
  • anon
    undefined
  • E-Reverance
    I really want to like ai music but none of them seem to be trained on reward models that reward "ambiance" or any sort of interesting sound design (which this paper obviously doesn't even concern), which might be reflective of the people training them not having niche music tastes
  • anon
    undefined
  • golem14
    apt reference:https://electricliterature.com/wp-content/uploads/2017/11/Tr..."The Petty and the Small; Are overcome with gall ; When Genius, having faltered, fails to fall. Klapaucius too, I ween, Will turn the deepest green To hear such flawless verse from Trurl's machine."
  • quaverquaver
    I guess it could be neat to play 100 of these at the same tempo at once but otherwise I'm having trouble coming up a use case for this...
  • rbtprograms
    all of these fall very flat, many of them sound like generic commercial tunes at their worst. there is obviously melody and structure but none of it sounds good.
  • halyconWays
    The music that YuE2 is makes is cleaner that Minimax Music 3, but I can't make it generate anything that isn't extremely bland. Songs fail to develop, they always sound like they're stuck in the intro. The composition feels extremely music-school-ey.
  • anon
    undefined
  • pdntspa
    Is there some way to get this to output just the vocals? Without running a stem splitter?It would be useful to be able to generate backing vocals/hums/melismas and whatnot, that could be plugged into a DAW.
  • hazenut
    As a musician, I have a feeling how this will go on:Most of the existing music creation heavily relies on daw and audio plugins, there are two major approaches among many: 1. recording with acoustic instruments, then using plugins to process 2.using a lot of virtual instruments directly, which is more often used in scoring.Some of the process is already heavily ai assisted, like the postprocessing and mastering. For scoring, more film scorers are using the Noteperformer which is directly note to audio but quality is far from Suno level. The very popular virtual singers(vocaloid, etc) evolved from audio concatenation to full neural network generations. But so far, none of these have encroached into the essential part of the music creation.Things will get more interesting if this get to the next level, when every single part of the music creation process (in the digital chain) is analyzable and synthesizable by ai. YuE2 appears to me can do note level analyze but stil not synthesize. To be able to synthesize, you get Noteperformer on steroid. Then, the daw and ai music creator will become one. This is somehow unappealing but paradoxical situation: You, as a creator, now can vastly do more. The ai tools gives you back the creative process. The ai music on the internet will no longer be pure slops because lots of them will be finely crafted by human musicians, but at the same time it all becomes super pointless because ai's capability is much stronger and do all these things itself, the only difference is whether you choose to do it yourself or let ai automate some of the parts.Therefore I'm not sure all the bad qualities of the ai music (slop, lacking soul, or whatever you name) is a sin or blessing. The condemnable part of the ai music is exactly what currently saves human musician.There are other side effects: the current virtual instrument, audio plugin market will mostly collapse. AI based audio generation will be a hegemony which reduces the diversity. You know that limitation of the process is a major source of curiosity and creativity, and itself is fun part. Wait until we rediscover the appeal of limitation, like the pixel art genre. We will see more of chiptunes, concatenative-sampling punks, 2010-digital vaporwaves and so on. But now the most important part is, genuineness will be a requirement: you have to show the process - you fake it with ai, you lose. Music will be a much more performative art (as it always was in the past) and community centric. Music streaming in its currently state of affair will be abandoned by all but the most casual listeners.Another trend will be gamifications and participatory music. There is still one part of music tech that is currently underexplored. The physical modeling. The synthesizer tech has hit a wall: it still struggle to generate natural and interesting sound, which is also why sample based virtual instrument still thrives. The physical modeling is promising but extremely compute intensive. AI accelerated physical modelling can help with this. This can be actually an antidote to the generative AI. Physical modeling can be enhanced by more visceral visualizations or hardware analogues (materialization of virtual process). There is joy of participating and watching the corporeal and physical process of making music. While elite artist can create very elaborate process(think of the eurorack and experimental synths) and visualizations and show them to the audience, the process can be simplified and gamified to promote collective creative performance, all without dumbing down, as the physical process itself is open ended, and we as human beings, have stornger connection to the physical sounds than the abstract sound of the previous gen synthesizers.To talk about participatory music and community, we can get glimpse of it from the Vocaloid scene. These crowd has very interesting take on ai music which I feel can be very enlightening to those uninitiated. They are enamored by the new AI soundbanks (which is the progression from the old contatenative technology) while very opposed to (generative) ai music. How is it so? They see the music by the Vocaloid producers a token of love poured into the community. The quality of music is not deciding factor, it is whether the producer know their community and pay respect to the virtual idol they loved. Their idol is a collective creation, it is not directed by any single entity. Derivative work also plays central role in the scene. Diaglogue is weighted a lot more than pure output. This is quite different from some of the AI idols though many uninitiated tend to conflate the two together. The AI industry may see Vocaloid as the predecessors to the AI idols, and you see some of the singing synth engine and soundbank makers are also giving nods to the AI music industry (ace studio, traditionally a singing synth company, is dipping toes in prompt based generations, and looks like is behind YuE2), but the two have very different spiritual cores. The conflation of AI music and Virtual Singers has caused further damage to the scene. One example is some streaming platform categorizes all virtual singer songs into AI music, which completely denies the hard work of the producers, and monetization and visibility possibilities, despite that some virtual singer still uses concatenative technology and have absolutely zero AI involved. But just for the AI soundbanks many has seen their drawbacks, now many consider they are too realistic. They want to keep the essence of the mechanical sound of their beloved idols, so you see the latest Miku soundbank delibrately not pursuing the realism like many do, e.g. SynthV, and those keep their idol unique characters are more successful in retaining their audience.
  • htx619
    [dead]
  • vivzkestrel
    [flagged]
  • dcw303
    I'm not species-ist. If a digital sentient being has something to say and they can express it through music, I'm more than willing to listen.But listening to these samples, it's obvious that, while technically proficient, they have no voice. Funnily enough, neither does most of the human produced music slop that existed before gen AI.All it exposes is that most people have no taste. This is as good as shuffling through popular songs on Spotify.