
[Keeping Tempo With Music Biz] — The Invisible Infrastructure AI Music Will Need: Op-Ed from Neutune

For the past few years, most conversations about AI music have focused on the models and what they can create. Then came the copyright lawsuits, questions about training data and, increasingly, licensing deals between technology companies and rightsholders.
There is another conversation sitting underneath all of this that we haven’t really started yet: if AI music becomes a significant part of the music economy, what infrastructure will actually make it work?
Music already relies on an enormous amount of infrastructure that listeners never see. Play a song and somewhere behind that three-minute stream are identifiers such as ISRC and ISWC, repertoire databases, DDEX messages, ownership data, royalty accounting, publishing administration, neighboring rights, collection societies, and systems for identifying content and reporting its use. All of this supports a transaction that is ultimately relatively simple to describe: a particular recording was played.
Generative music creates a very different level of complexity. A recording might be used while only its melody remains in the resulting output. The vocals could be transformed, a drum component could come from another source, the user could add an instrument, the AI could generate new material, eight bars could be regenerated, or another licensed component could be introduced later. The resulting music itself may even become the starting point for another generation.
The infrastructure therefore has to answer questions it wasn’t built to answer. What was used, which part was used, where did it come from, was it licensed, how was it transformed, and how much survived into the output? It also needs to connect that information to the relevant rights, attribution and ultimately remuneration, all while preserving enough of the history for it to be understood once the music leaves the environment where it was created.
The music industry built infrastructure to identify songs. AI will require infrastructure that can identify what is happening inside of songs.
A new infrastructure stack
There won’t be one database, identifier or standard capable of answering all of these questions. Several layers will have to work together.
Identity comes first because we need to know what the content is and what components exist within it. Provenance establishes where those components came from and what happened to them, while attribution connects them to existing works, recordings, performances, vocals, lyrics or other protected material. Rights information establishes who controls those elements and the permissions attached to them.
That information then has to connect with actual usage during generation or transformation and, eventually, with the economics of that usage. There also needs to be a way for rightsholders, platforms and other participants to verify what happened, particularly when information is missing, and for all of this data to move between AI companies, DSPs, labels, publishers, distributors and rights organizations.
The music industry already has pieces of this puzzle. ISRC and ISWC provide identity at the recording and composition levels. DDEX allows enormous amounts of information to move between companies in standardized formats. Rights databases connect repertoire to ownership. Fingerprinting and other recognition technologies help identify existing recordings. C2PA is developing a broader technical framework for content provenance and authenticity.
AI is arriving in an industry with decades of infrastructure behind it, but that infrastructure was designed around different types of transactions. One of the biggest differences is the unit of music that needs to be understood.
What happens when the song stops being the smallest useful unit?
The music business has spent decades organizing data and rights around relatively stable objects such as a composition, a recording, an album or a stream. Generative AI can operate several layers below them.
A melody can remain recognizable while almost everything surrounding it changes. A vocal can be separated from a recording and transformed. A musical segment can contribute to an output without the complete source recording appearing in it. Components associated with several works or recordings can coexist in a single generation.
An ISRC can tell us which recording we are talking about, and an ISWC can tell us which musical work. But neither was designed to identify an individual piece within them.
That has consequences far beyond identification. If the infrastructure cannot persistently identify a musical component, maintaining provenance for it becomes difficult, as does connecting it to rights information, recording its use during generation, or following what happens when it is transformed.
AI therefore creates the need for another identification layer below the recording and composition, which immediately raises another problem: knowing where those pieces came from and what has happened to them.
Provenance becomes much harder when music is fluid
Provenance sounds relatively straightforward when applied to a complete digital object. We can ask where a file originated, who created it, whether AI was involved and what happened to it along the way.
Generative music complicates that idea because the history of the file and the history of the music inside the file aren’t necessarily the same thing. A track can be genuinely new as a file while containing musical elements connected to existing material. It can combine human-created, licensed & generated components, some of which may survive through several transformations while others disappear after one generation.
Knowing that a track is “AI-generated” therefore tells us relatively little about those relationships. Attribution becomes part of the provenance problem because we also need to understand the origins of the musical elements themselves.
There are two moments when that attribution can potentially be established. If a licensed AI system knows which source material is being used, it can record those contributions as the music is generated, so the provenance history begins with the creation itself. In other situations, a track will arrive without a reliable generation history and will need to be analyzed to establish whether identifiable existing material contributed to it.
These are two very different technical problems: attribution by design and attribution after the fact. One records what happens during generation; the other attempts to reconstruct that history from the resulting music. Both will be necessary, because even excellent provenance systems won’t result in every piece of music arriving everywhere with a complete history attached.
What would this look like in practice?
This is precisely the problem we have been exploring at Neutune, an AI research lab and music technology company developing technologies for music creation, attribution and rights management. Based in South Korea and working globally across both music and tech, Neutune combines AI research with practical applications designed to enable new ways to create, transform, understand and track music while maintaining the connection between creative content, its origins and its rights.
Its technologies include MixAudio, a licensed AI music creation & remix platform with attribution built directly into the generation process, tracking the components used to create each output and their associated rights; and musicDNA, a B2B technology for attribution, provenance, copyright infringement detection and AI-generated music detection. musicDNA analyzes music at the component and segment level to identify where musical material comes from, detect similarities and potential unauthorized uses across works, trace how content has been transformed or reused, and determine whether music or individual components show evidence of AI generation.
Working on both has given us a particular view of the infrastructure problem because we encounter it from both directions. Within a generation environment, we can know what is happening as music is created and transformed. With finished audio, we face the opposite problem and have to determine what happened before the music reached us.
This led us to think about music in terms of atomic musical components. With musicDNA, a piece of music can be analyzed at component level, looking at relationships across the musical work, sound recording, vocals and lyrics. The relevant question becomes more granular than simply identifying a song. We want to understand which musical elements are present, where they originated and what happened to them.
When taking a deeper look, the identity problem immediately becomes apparent. A recording has an ISRC and a composition has an ISWC, but there is no equivalent widely adopted identifier for the individual component being attributed. That is the reason we developed the International Standard Block Code, or ISBC, with the idea of creating an identifier at component level that can connect back to the recording and composition layers already used by the industry.
Inside MixAudio, the source environment is known. The platform works with licensed music and users can transform elements of it, changing arrangement, instrumentation, tempo, timbre and structure, adding licensed components or their own vocals and instruments. Since the system knows what enters the generation, component usage can be recorded while the music is being created. This gives us a practical environment for attribution by design.
musicDNA addresses the other direction. When the generation history isn’t available, the resulting audio can be analyzed to establish relationships with reference material at component level. And BlockDB explores another part of the same problem by connecting components with provenance, rights information and usage events.
These technologies are our attempt to understand what infrastructure built around components could look like in practice. They have also made clear how much larger the problem is than any individual technology. An identifier has limited industry value if other systems cannot understand it. Provenance loses much of its value if it disappears when music leaves the platform where it was created. Attribution cannot support remuneration unless it connects to rights and usage information. The difficult part is getting those layers to work together.
Licensing will expose the next gap
The music industry’s AI licensing conversation is currently dominated by permission and price: which catalogues can be used, for what purposes, under which conditions and how much AI companies should pay.
If the industry makes substantial progress on those questions, the next problem appears very quickly. An AI company could have licensed access to millions of recordings and compositions, while its users create millions of outputs from that repertoire. The licence establishes what the company is allowed to use, but the economic system underneath the licence still needs some way of determining what was actually used.
Generative usage can be considerably more complex than a stream. Several components may contribute to one output, a component may disappear during regeneration, another may remain through successive versions, and different rights may attach to the recording, composition, vocal or other material involved.
This is where attribution starts becoming part of the accounting infrastructure. Granular attribution can provide information about what contributed to a generation, usage data can record when those contributions occurred, and rights information can connect them to the relevant parties. Together, those layers can provide the basis for different forms of remuneration.
The eventual economic model is a separate question. The industry could use pools, per-generation models, component-based allocation, minimum guarantees, any combination of these models, or even approaches we haven’t developed yet. The underlying infrastructure needs to make the activity measurable enough for those different models to be possible.
We keep asking whether AI companies will license music. The next question is harder: once they do, what infrastructure will tell us what was actually used and who should be paid?
Provenance that cannot travel isn’t enough
Even if identity and attribution work extremely well within individual platforms, music doesn’t remain inside a single system. It can leave an AI service, go through a distributor, arrive on a DSP, appear in a video, be edited and eventually enter another creative environment.
A detailed provenance record sitting in the database of the original AI company doesn’t automatically travel with it, which makes interoperability one of the least visible and potentially most consequential parts of AI music infrastructure.
The industry already understands the principle. DDEX exists because digital music requires different companies to exchange information in ways their systems can understand. AI will add new information to that chain, including AI involvement, component identity, provenance, attribution, transformations and potentially component-level usage.
Broader initiatives such as C2PA can provide part of the provenance layer, while music also has requirements created by its particular rights structures and by the relationships between compositions, recordings, performances and the components within them.
The infrastructure also has to accommodate imperfect information. Metadata gets lost, files are transformed and re-uploaded, generation histories aren’t always available, and some systems will provide far more detailed provenance than others. This is why verification has to sit alongside provenance. Sometimes the infrastructure will carry the history of a piece of music, and sometimes technology will have to investigate the content itself to reconstruct part of that history.
The scale problem
Streaming works because the industry doesn’t manually investigate every stream. Infrastructure turns an enormous number of individual events into data that computers can process, exchange and eventually settle.
Generative music could produce an even larger number of events, with each one potentially much harder to describe. If every AI-generated track requires lawyers, musicologists and forensic experts to reconstruct which material was used, where it came from, how it changed and who should be compensated, the economics won’t scale.
The infrastructure underneath AI music will therefore need ways to identify music at the level where AI operates, preserve provenance as musical elements move through generations, connect attribution with rights and usage, verify content when the history is incomplete, and allow that information to move between companies.
At Neutune, we are working on one part of this much larger problem, particularly around component-level identity, attribution and provenance. Other companies, rightsholders, DSPs, standards organizations and researchers will solve other parts. Some of the infrastructure will evolve from systems the music industry already has, while other pieces still need to be built.
If the industry cannot identify those components and preserve their history, the consequences will follow through the entire chain. A rightsholder cannot reliably know where its repertoire contributed to generated music. A licensing partner cannot report precisely how licensed material was used. Provenance can disappear when content moves between systems. Attribution cannot reliably connect to remuneration. And a platform receiving the finished track may have no way to establish what it contains or where its components came from.
This is where the scale of the problem becomes clear. One generation can involve several works, recordings and rightsholders, as well as multiple transformations and successive versions. Multiply that by millions of users generating and modifying music continuously, and the industry cannot rely on metadata attached to a finished track or on investigations after distribution. The identity and history of the music have to be created, preserved and communicated as part of the transaction itself, with ways to reconstruct that history when it is missing.
That is what will determine whether millions of AI music transactions can actually be attributed, licensed, reported, audited and paid for. Without it, we may have an AI music market, but we won’t have the infrastructure to know what is happening inside it.
Written by: Dr. Jongpil Lee, Co-Founder & CEO, and Virginie Berger, Chief Industry & Rights Officer, Neutune
You can read past Keeping Tempo’ articles via the portal linked here. And, stay tuned for more insightful discussions from our members and partners from across the industry!










