Are there AI models for generating sounds based on a text and reference?

I've been having a hard time finding a solution. Is there really no commercialized model that I can feed in a reference sound and text instruction and get another sound out?Right now having a multimodal inputs to image or text output is a commodotized, solved problem. IE you can put a prompt for some image and use a reference image to guide the model on what you want. After all, a picture is worth a thousand words right?Ive also used a text and audio input in order to get a text description or classification out.I cannot for the life of me find a solution for Audio + text -> AudioMy usecase is that I'm trying to generate new sound effects based on a source that doesn't have that many clean examples and thought this would be just another solved workflow but I'm not finding much. Tried elevenlabs SFX but it's only text->audio and without the reference, it's hard to guide with just text and recreate a sound with any accuracy. I tried the Stable Audio 3 model which is supposed to do what I want but the results were terrible. This modality seems to be stuck where image generation was in 2016. Is there something else or is this just not a common need?

Hacker NewsToolAISource
0Sign in to voteCopy link

FL score

70

out of 100

Verdict

VALIDATE

high confidence

Competition

No competitor data yet

Trend

No signal yet

A tool to generate new sound effects from text instructions and reference audio to fill a gap in current AI audio generation capabilities.

The pain

Creators struggle to generate accurate, guided sound effects because existing models either only use text or produce poor quality when combining text and audio inputs.

The gap

No widely available commercial AI model effectively combines text and reference audio to produce high-quality, controllable new sounds.

Build angle

Focus on improving multimodal audio generation models with better training data and user controls, possibly leveraging advances in related AI fields like image generation.

Strengths

  • Clear and specific unmet need in audio generation
  • Existing AI models have not solved this, indicating opportunity
  • Potential to serve sound designers, game developers, and media producers

Risks

  • Technical complexity of multimodal audio generation
  • Uncertain market size and willingness to pay
  • Competition from incremental improvements in text-to-audio models

Questions about this idea?

FlyBot reads the scoring and gives you a second opinion on “Are there AI models for generating sounds based on a text and reference?”.

Open FlyBot