- Loop Showcase
- One-Shot Showcase
- Keybed Showcase
- Multi-Layered Keybeds
- 1. Musical Structure
- 2. Instrument Identity
- 3. Timbral Control
- 4. Timbral Mixing
- 5. FX Prompting
- 6. Loop Fidelity
- 7. One-Shot Generation
- 8. Pitch-Consistent Keybeds
- Hierarchy Overview
- Major Families
- Sub-Family Coverage
- Keys and Modes
- Loop Structure
- Layered Prompt Structure
- Prompting Notes
- Loop Generation
- Samples / One-Shot Generation
- Keybed / Text-to-Synth Generation
- Recommended Interfaces
- Model Files
- Basic Setup for RC Stable Audio Tools
- Hardware Requirements
- Generation Performance
- Specialized Samples and Keybeds Training
- Keybed-Specific Limitations
- Foundation-1 Overview
- Foundation-1.2 Samples & Keybeds Update
- Keybed Guided Demo
Foundation-1
Structured text-to-sample and text-to-instrument generation for modern music production
Overview
Foundation-1 is a family of text-conditioned audio models designed around music production workflows. Rather than treating audio generation as broad caption-to-music generation, Foundation-1 was trained around structured controls for instrument identity, timbre, FX, musical behavior, pitch, timing, and tonality.
The original Foundation-1 checkpoint focuses on structured loop generation. Newer specialized checkpoints extend the same conditioning system into individual one-shots and pitch-consistent playable keybeds.
This means Foundation-1 can now be used at several levels of a production workflow:
- generate a tempo-synced musical loop
- generate an individual note or sound
- generate multiple pitch-consistent notes from the same text prompt
- assemble those notes into a playable sampler instrument
- combine multiple generated keybeds into layered instruments
The goal is not simply to generate finished audio. The specialized keybed workflow is designed to let the model generate the source material for an instrument, then hand control back to the producer.
Foundation-1 Model Family
| Model | Primary Use | Notes |
|---|---|---|
| Foundation-1 | Loop generation | Original checkpoint for structured, BPM-aware, bar-aware musical loops |
| Foundation-1.2 Samples | Loop Generation / One-Shot generation | Selected earlier in specialized training to preserve stronger short-form loop-generation quality |
| Foundation-1.2 Keybeds | One-Shot generation, Pitch-consistent sample generation and playable instruments | Trained longer to reinforce timbral consistency across pitch and register |
The Foundation-1.2 Keybeds model is not limited to full keybed generation. It remains highly capable at generating individual samples and one-shots. Its specialization reflects the additional training needed to maintain a consistent sonic identity across many generated pitches.
The checkpoints were separated because the longer keybed-focused training improved cross-pitch consistency, while reducing the broader loop-generation quality preserved by the earlier Samples checkpoint.
Text-to-Synth / Keybed Generation
The Keybed checkpoint extends Foundation-1 from text-to-sample generation into a practical text-to-instrument workflow.
A user can describe a sound using the same instrument and timbre vocabulary used elsewhere in Foundation-1, generate pitch-specific samples across a keyboard, and automatically assemble the result into a playable sampler instrument.
Examples can range from conventional instruments:
Grand Piano, Warm, Gritty, Wet, Low Reverb
to synthetic or hybrid sounds:
Cello, Harp, Rich, Clean, Choir, Pluck, Formant Vocal, Wet, High Reverb
Because pitch is generated rather than conventionally pitch-shifted from a single root sample, the model can create natural timbral variation across registers while maintaining the identity of the prompted sound.
For the intended workflow, the Keybed model is best used with RC Stable Audio Tools, which handles the multi-note generation process and instrument assembly automatically. Generated keybeds can be exported for use as playable sampler instruments rather than remaining as isolated audio generations.
The user-facing prompt is intentionally simple. A descriptor such as:
Grand Piano, Warm, Gritty
is expanded internally into the structured Keybed conditioning grammar used by the model. RC Stable Audio Tools adds the required Keybed / sequence / note information, keeps the user descriptor stable across the instrument, reuses the same resolved seed across generation chunks, then slices and maps the resulting notes into the exported sampler.
This makes the workflow useful not only as a demo interface, but also as a reference implementation for developers interested in building their own VST, sampler, or instrument-generation front end around the model.
For the full prompt-injection and inference design, see the Keybed Training & Inference Strategy.
Foundation-1.2 Keybed generation workflow in RC Stable Audio Tools.
What Foundation-1 Does
- Generates musically coherent loops for production workflows
- Generates individual note-specific one-shots
- Generates pitch-consistent multi-note keybeds
- Builds playable sampler instruments through the RC Stable Audio Tools workflow
- Supports layered keybeds built from multiple independently prompted sounds
- Understands BPM and bar count for structured loop generation
- Locks to major and minor keys across western music theory
- Supports enharmonic equivalents when prompting scales and keys
- Separates instrument identity from timbral character
- Supports timbral mixing by combining instrument and sonic descriptors
- Responds to FX tags such as reverb, delay, distortion, and modulation
- Uses notation-style prompt structure to encourage coherent phrasing, melodic shape, rhythmic behavior, and harmonic motion
- Produces perfect loops within supported BPM / bar denominations
- Understands Wet vs Dry production context β adding terms like Dry encourages minimal FX processing, while Wet or FX tags produce more processed, spatial, or effected sounds.
Why It Feels Different
Most audio models can react to broad prompt terms like βwarm padβ or βbright synth.β with inconsistent results. Foundation-1 was designed to go further by treating the sound as a layered system:
- Instrument Family β what broad source category the sound belongs to
- Sub-Family β the more specific instrument role or identity
- Timbre Tags β the tonal, spectral, or textural character
- FX Tags β the processing layer applied to the sound
- Notation / Structure Tags β the musical behavior of the generated phrase
This layered conditioning approach is a major reason Foundation-1 is able to deliver both high musicality and high prompt control at the same time.
Audio Showcase
Loop Showcase
| Prompt | Audio |
|---|---|
| Bass, FM Bass, Medium Delay, Medium Reverb, Low Distortion, Phaser, Sub Bass, Bass, Upper Mids, Acid, Gritty, Wide, Dubstep, Thick, Silky, Warm, Rich, Overdriven, Crisp, Deep, Clean, Pitch Bend, 303, 8 Bars, 140 BPM, E minor | |
| Sub Bass, Bass, Gritty, Small, Square, Bass, Dark, Digital, Thick, Clean, Simple, Bassline, Epic, Choppy, Melody, 4 Bars, 150 BPM, G# minor | |
| Flute, Pizzicato, Punchy, Present, Ambient, Nasal, Melody, Epic, Airy, Slow Speed, 8 Bars, 150 BPM, E minor | |
| High Saw, Spacey, Lead, Warm, Silky, Smooth, 303, Synth Lead, Medium Reverb, Low Distortion, Upper Mids, Mids, Pitch Bend, Arp, 8 Bars, 140 BPM, F minor | |
| Trumpet, Warm, Complex Arp Melody, High Reverb, Low Distortion, Smooth, Silky, Texture, 8 Bars, 130 BPM, C minor | |
| Synth, Pad, Chord Progression, Rising, Digital, Bass, Fat, Near, Wide, Silky, Warm, Focused, 8 Bars, 110 BPM, D major | |
| Piccolo, Flute, Airy, Music Box, plucked, complex melody, 8 Bars, 140 BPM, C# minor | |
| Synth Lead, Wavetable Bass, Low Distortion, High Reverb, Sub Bass, Upper Mids, Acid, Gritty, Wide, Thick, Silky, Warm, Rich, Overdriven, Crisp, Clean, 303, Complex, 8 Bars, 140 BPM, F minor | |
| Fiddle, Bowed Strings, Full, Clean, Spacey, Rich, Intimate, Thick, Rolling, Arp, Fast Speed, Complex, 8 Bars, 128 BPM, B minor | |
| Chiptune, Chord Progression, Pulse Wave, Medium Reverb, 8 Bars, 128 BPM, D minor | |
| Kalimba, Mallet, Medium Reverb, Overdriven, Wide, Metallic, Thick, Sparkly, Upper Mids, Bright, Airy, Alternating, Chord Progression, Atmosphere, Spacey, Fast Speed, 8 Bars, 120 BPM, B minor |
One-Shot Showcase
The examples below show short, note-specific one-shot generations across acoustic, synthetic, bass, FX, and hybrid timbres.
| Prompt | Note | Audio |
|---|---|---|
| FX, Sharp, Subdued, Distant, Sparkly, Round, Hit, Deep, Low Reverb, Sub | D#1 | |
| Synth Bass, Metallic, Rich, Punchy, Digital, Medium Reverb | C2 | |
| Reese Bass, Hit, Bright, Wavetable, Noisy, Big, Thick, Airy, Deep | F#2 | |
| Synth Lead, Shiny, Rich, Choir, Bright, Smooth, Short, Synthetic Vox | F5 | |
| Tuba, Big, Bright, Clean, Biting, Sustained | G#3 | |
| Clavinet, Wobble, Breathy, Bright, Deep, Gritty, Impact, Sparkly | C#4 | |
| Cello, Airy, Wide, Woody, Muffled, Smooth | D#3 | |
| Reese Bass, Subdued, Round, Growl, Hit, Deep, Woody, Fat, Metallic, Medium Delay, Medium Distortion | D#2 | |
| Digital Piano, Retro, Smooth, Mono, Warm, Low Reverb | E3 | |
| Supersaw, Big, Smooth, Warm, Vintage, Crisp, Analog, Wide, Muffled, High Reverb | G2 | |
| Synth Lead, Bright, Shiny, Soft, Warm, Punchy, Smooth, Medium Delay | C#4 | |
| Sustained, Trumpet, Airy, Wide, High Reverb | F3 | |
| 808, Thick, Acid, Hit, FX, Rumble, Wide, Big, Distortion | F1 | |
| Violin, Vintage, Thick, Chiptune, Spacey, Nasal, Medium Delay | A#5 | |
| Grand Piano, Near, Snappy, Distant, Bell, Bright, Focused, Delay | D3 |
Keybed Showcase
The following examples demonstrate generated keybeds across acoustic, synthetic, and hybrid timbres.
A full video playthrough of the Keybed Showcase is available here, which shows the associated MIDI used to audition each generated instrument.
| Prompt | Audio |
|---|---|
| Soft, Silky, Spectral, Smooth, Spacey, Synth, Pad, Subdued, Wet, High Reverb | |
| Pad, Focused, Buzzy, Big, Deep, Supersaw, Pulse, Dry | |
| Pan Flute, Hollow, Dark, Smooth, Sustained, Wide, Silky, Wet, Medium Reverb | |
| Violin, Woody, Pizzicato, Focused, Breathy, Airy, Wet, High Reverb | |
| Grand Piano, Warm, Gritty, Wet, Low Reverb | |
| Grand Piano, Cold, Sparkly, Wet, Low Reverb | |
| Digital Piano, Noisy, Spiccato, Rich, Pluck, Warm, Swell, Clean, Dry | |
| Cello, Fat, Warm, Metallic, Sharp, Near, Sustained, Dry | |
| Church Bell, Glassy, Sparkly, Wet, High Reverb | |
| Thick, Saw, Neuro, Reese Bass, Dry | |
| Thick, Sine, Reese Bass, Wet, High Reverb, High Phaser | |
| Xylophone, Sustained, Analog, Warm, Woody, Dry | |
| Xylophone, Sustained, Analog, Warm, Woody, Bit Crushed, Wet | |
| Cello, Harp, Rich, Clean, Choir, Pluck, Formant Vocal, Wet, High Reverb | |
| Acid, Neuro, Synth Lead, Saw, Square, Wide, Focused, Pluck, Sustained, Wet, High Phaser | |
| Synth Lead, Present, Sharp, Spacey, Digital, Hollow, Focused, Clean, Wet, Cross Delay, High Reverb | |
| Dubstep, Neuro, Synth Lead, Pitch Bend, Acid, Wet, Cross Delay |
Generated Waveforms
Pure waveform keybeds are shown separately to make the learned oscillator shapes easier to compare directly.
| Waveform | Prompt | Audio |
|---|---|---|
| Triangle | Pure Tone, Triangle, Dry | |
| Square | Pure Tone, Square, Dry | |
| Pulse | Pure Tone, Pulse, Dry | |
| Sine | Pure Tone, Sine, Dry | |
| Saw | Pure Tone, Saw, Dry |
Multi-Layered Keybeds
The inference pipeline can also combine three independently generated keybeds into a single layered instrument. The examples below use separate Main and Support prompts for each layer. Further the generated instrument allows independent volume control for each layer.
| Layer Prompts | Audio |
|---|---|
| Main: Bright, Airy, Violin Support 1: Thick, Present, Male, Vocal, Choir Support 2: Digital String, Analog |
|
| Main: Marimba, Dark, Sub, Woody, Airy Support 1: Soft, Grand Piano, Dark Support 2: Music Box, Glassy |
Core Capabilities
1. Musical Structure
Foundation-1 was trained to produce structured musical material rather than full music or generic textures. Musical Notation terms can encourage notation, chord progressions, melodies, arps, phrase direction, rhythmic density, and other musically relevant behaviors.
2. Instrument Identity
The model supports a broad instrument hierarchy spanning synths, keys, basses, bowed strings, mallets, winds, guitars, brass, vocals, and plucked strings.
3. Timbral Control
Foundation-1 is not limited to broad instrument naming. It also responds to timbral descriptors such as spectral shape, tone, width, density, texture, brightness, warmth, grit, space, and other sonic traits.
4. Timbral Mixing
Because instrument identity and timbral character were not collapsed into a single flat label, the model is especially strong at timbral hybridization and layered sonic prompting.
5. FX Prompting
The model supports a dedicated FX layer covering multiple forms of reverb, delay, distortion, phaser, and bitcrushing.
6. Loop Fidelity
Foundation-1 is built for production-ready loop generation, including BPM-aware and bar-aware structure within supported denominations.
7. One-Shot Generation
The specialized Foundation-1.2 Samples and Keybeds checkpoints generate short, note-specific sounds across acoustic, synthetic, bass, FX, and hybrid timbres. These can be used individually or as raw material for sampler instruments.
8. Pitch-Consistent Keybeds
The Keybed checkpoint was trained specifically to preserve a prompted timbral identity across changing pitch and register. RC Stable Audio Tools uses this capability to generate multiple note-specific samples and assemble them into playable instruments.
Conditioning Architecture
Foundation-1 was trained with a layered tagging hierarchy designed to improve control, composability, and prompt clarity.
Hierarchy Overview
- Major Family β broad instrument class
- Sub-Family β more specific instrument role
- Timbre Tags β tonal / spectral / textural descriptors
- FX Tags β processing layer
- Notation Tags β musical behavior and phrasing
This makes it possible to prompt at different levels of abstraction. A user can stay broad with a family-level prompt like Synth or Keys, or get more specific with terms like Synth Lead, Wavetable Bass, Grand Piano, Violin, or Trumpet, then further shape the output using timbral and FX descriptors.
Instrument Coverage
Major Families
Foundation-1 was trained across the following major instrument families:
- Synth
- Keys
- Bass
- Bowed Strings
- Mallet
- Wind
- Guitar
- Brass
- Vocal
- Plucked Strings
Sub-Family Coverage
Foundation-1 includes a wide sub-family layer covering a broad range of production-relevant instrument roles, including but not limited to:
- Synth Lead
- Synth Bass
- Digital Piano
- Pluck
- Grand Piano
- Bell
- Pad
- Atmosphere
- Digital Strings
- FM Synth
- Violin
- Digital Organ
- Supersaw
- Wavetable Bass
- Rhodes Piano
- Cello
- Texture
- Flute
- Reese Bass
- Wavetable Synth
- Electric Bass
- Marimba
- Trumpet
- Pan Flute
- Choir
- Harp
- Church Organ
- Acoustic Guitar
- Hammond Organ
- Celesta
- Vibraphone
- Glockenspiel
- Ocarina
- Clarinet
- French Horn
- Tuba
- Oboe
Timbre System
One of Foundation-1βs main strengths is that it was not trained to treat timbre as an afterthought. Timbral character is directly represented in the prompt system, giving users control over not only what is being generated, but also how it sounds.
Representative timbre descriptors include:
- Warm
- Bright
- Wide
- Airy
- Thick
- Rich
- Tight
- Full
- Gritty
- Clean
- Retro
- Saw
- Crisp
- Focused
- Metallic
- Chiptune
- Dark
- 303
- Shiny
- Analog
- Present
- Sparkly
- Ambient
- Soft
- Smooth
- Cold
- Buzzy
- Deep
- Formant Vocal
- Round
- Punchy
- Nasal
- Vintage
- Growl
- Breathy
- Glassy
- Noisy
- Synthetic Vox
- Supersaw
- Bitcrushed
- Dreamy
Why This Matters
This tagging design makes prompts much more flexible. Instead of only asking for an instrument, users can shape:
- tonal balance
- brightness / darkness
- width / intimacy
- clean vs driven character
- synthetic vs organic feel
- transient sharpness
- texture and density
- spatial character
This is especially useful for producers who want to guide the output toward a specific role in a mix rather than just a generic instrument label.
For a list of used tags please see the Tag Reference Sheet.
FX Layer
Foundation-1 includes a dedicated FX descriptor layer spanning multiple common production effects.
Representative FX tags include:
- Low Reverb
- Medium Reverb
- High Reverb
- Plate Reverb
- Low Delay
- Medium Delay
- High Delay
- Ping Pong Delay
- Stereo Delay
- Cross Delay
- Mono Delay
- Low Distortion
- Medium Distortion
- High Distortion
- Phaser
- Low Phaser
- Medium Phaser
- High Phaser
- Bitcrush
- High Bitcrush
Musical Notation and Structure
Foundation-1 was trained with structured musical descriptors designed to improve phrase coherence, rhythmic intent, melodic motion, and prompt control.
These notation-style prompt terms help steer:
- chord progressions
- melodies
- top-line layers
- arpeggios
- phrase direction
- rhythmic density
- harmonic feel
- subdivision style
- simple vs complex motion
- sustained vs plucked behavior
- melodic contour and pacing
Examples of supported structural ideas may include terms such as:
- chord progression
- melody
- top melody
- arp
- triplets
- simple
- complex
- rising
- falling
- strummed
- sustained
- catchy
- epic
- slow
- fast
This notation layer is one of the main reasons Foundation-1 produces unusually coherent musical material instead of static or loosely related phrases. These can be mixed and matched as desired.
Tonal and Timing Support
Foundation-1 is designed for structured music production workflows and supports:
Keys and Modes
- Major keys
- Minor keys
- Enharmonic equivalents
- Western 12-tone chromatic prompting
Loop Structure
- Supported bar lengths: 4 Bars, 8 Bars
- Supported BPM denominations: 100 BPM, 110 BPM, 120 BPM, 128 BPM, 130 BPM, 140 BPM, 150 BPM
Prompt Structure
For best results, use rich prompts built around the modelβs tags. These tags can be mixed and matched as needed. The model was trained on a structured hierarchy designed to encourage musically coherent sample generation.
Layered Prompt Structure
[Instrument Family / Sub-Family], [Timbre], [Musical Behavior / Notation], [FX], [Key], [Bars], [BPM]
Prompting Notes
- Start with a clear instrument identity
- Add 1β3 timbre descriptors for stronger steering
- Include a notation or musical structure term for better phrase coherence
- Always include Bars and BPM, which define the musical loop length
- Ensure the generation duration matches the requested musical structure
- The RC Stable Audio Fork automatically handles this timing alignment
Use FX and timbre tags sparingly at first, then layer more once you understand the modelβs behavior.
One Prompt β Multiple Outputs
Each row below uses the exact same prompt, but a different random seed.
The timbre tags remain unchanged, so the overall sound character stays consistent while the melodic and musical content varies between generations.
| Prompt | Output A | Output B | Output C |
|---|---|---|---|
| Bass, FM Bass, Medium Delay, Medium Reverb, Low Distortion, Phaser, Acid, Gritty, Wide, Dubstep, Thick, Silky, Warm, Rich, Overdriven, Crisp, Deep, Clean, Triplets, 8 Bars, 150 BPM, A minor | |||
| Gritty, Acid, Bassline, 303, Synth Lead, FM, Sub, Upper Mids, High Phaser, High Reverb, Pitch Bend, 8 Bars, 140 BPM, E minor | |||
| Kalimba, Mallet, Medium Reverb, Overdriven, Wide, Metallic, Thick, Sparkly, Upper Mids, Bright, Airy, Small, Alternating Chord Progression, Atmosphere, Spacey, Fast, 4 Bars, 120 BPM, B minor |
Recommended Workflow
Foundation-1 is best used with RC Stable Audio Tools, which is tuned around the model family, its metadata, and its structured prompting system.
RC Stable Audio Tools (Enhanced Fork)
The interface supports different workflows depending on the checkpoint in use.
Loop Generation
For the original Foundation-1 loop checkpoint, the interface provides:
- structured prompt building aligned with the training tags
- random prompt generation
- automatic BPM / bar timing alignment
- automatic MIDI extraction from generated audio
- generation settings tuned for Foundation-1
Samples / One-Shot Generation
The specialized Samples workflow supports:
- note-specific generation
- instrument and timbre prompting
- short sample-focused inference
- rapid auditioning of generated sounds
The Keybed checkpoint can also generate individual one-shots effectively. The dedicated Samples checkpoint is provided because its earlier training endpoint preserves stronger general sample-generation quality.
Keybed / Text-to-Synth Generation
The Keybed workflow is designed to be used in conjunction with RC Stable Audio Tools rather than as a sequence of manually generated independent notes.
The interface handles the pitch-aware generation pipeline, creates the source files required for the instrument, and can export completed keybeds for sampler use.
The workflow supports:
- automatic generation across keyboard registers
- pitch-specific text-conditioned samples
- playable keybed construction
- DecentSampler export
- SFZ export
- per-instrument ADSR controls
- built-in sampler effects and tone shaping
- layered keybeds using independently prompted Main / Support layers
- independent layer volume control
- generated instrument previews
Generated keybeds can be assembled and exported as playable DecentSampler or SFZ instruments.
Recommended Interfaces
RC Stable Audio Tools (Enhanced Fork)
Stable Audio Tools (Original Repository)
Model Files
Foundation-1 is now distributed as multiple specialized checkpoints built around the same underlying architecture and conditioning system.
The specialized Foundation-1.2 Samples and Keybeds checkpoints retain the same Foundation-1 / SAO architecture. They therefore do not represent separate model architectures; the specialization comes from the training objective and selected checkpoint.
The release uses 16-bit model weights to reduce the model footprint without changing the intended inference quality.
Foundation_1.safetensorsβ loop-generation checkpointFoundation-1.2-Samples.safetensorsβ sample-focused checkpointFoundation-1.2-Keybeds.safetensorsβ keybed-focused checkpointmodel_config.jsonβ shared model configuration
Basic Setup for RC Stable Audio Tools
- Create a subfolder inside your
modelsdirectory - Place the desired checkpoint and its compatible config in that folder
- Launch the interface
- Select the checkpoint from the model selector
- Choose the matching Loop, One-Shot, or Keybed workflow
- Prompt with layered instrument and timbre descriptors
For full keybeds, the RC interface is strongly recommended because it manages the multi-note inference and sampler-export process automatically.
Hardware Requirements
Foundation-1 is designed to run locally on modern GPUs.
Typical VRAM usage during generation is approximately ~7 GB.
For reliable operation, a GPU with at least 8 GB of VRAM is recommended.
Generation Performance
Generation speed will vary depending on GPU model, generation mode, sample length, and system configuration.
On an RTX 3090, a standard individual generation is approximately ~7β8 seconds per sample. Full single-layer keybed builds take approximately 50 seconds.
Dataset and Training Philosophy
Foundation-1 was built around a structured sample-generation philosophy, rather than generic or genre-based audio captioning. The dataset consists entirely of hand-crafted and labeled audio, produced through a controlled augmentation pipeline.
At a high level, the training design emphasizes:
- structured musical loops
- instrument hierarchy
- explicit timbre representation
- dedicated FX descriptors
- notation-aware prompt terms
- strong production relevance
- broad reuse for compositional workflows
This design is central to the modelβs musical coherence and high degree of sonic control.
For more details on the original dataset and training methodology, see the Training & Dataset Notes.
Specialized Samples and Keybeds Training
The Foundation-1.2 Samples and Keybeds checkpoints use the same underlying Foundation-1 architecture, but were trained with a lower learning rate of 1e-5.
| Checkpoint | Training Endpoint | Specialization |
|---|---|---|
| Foundation-1.2 Samples | Epoch 0 / Step 560 | General sample / one-shot generation |
| Foundation-1.2 Keybeds | Epoch 3 / Step 8120 | Cross-pitch timbral consistency and playable keybed generation |
Additional training configuration:
- Optimizer: AdamW
- Learning Rate:
1e-5 - Weight Decay:
1e-3 - Scheduler: InverseLR
- EMA: Enabled
- Sample Rate: 44,100 Hz
- Channels: Stereo
- Training Window: 882,000 samples (~20 seconds)
The longer Keybed run was selected to reinforce cross-note and cross-register consistency. The objective was not simply to improve isolated note quality, but to teach the model how a single sound identity should behave as pitch changes across an instrument.
The Keybed checkpoint used a dedicated multi-note training strategy rather than treating every pitch as an unrelated example. The generation pipeline mirrors that structure by injecting note-sequence conditioning, reusing a common seed across keybed chunks, slicing generated sequences back into individual notes, and packaging the resulting samples into playable instruments.
For a detailed explanation of both the training method and the inference/export pipeline, see the Keybed Training & Inference Strategy.
Limitations
Foundation-1 is a specialized model family for producer-facing sample and instrument generation, not a general-purpose full-song generator.
Important notes:
- It performs best when prompted using vocabulary aligned with the training design
- It is optimized for sample-generation workflows, not open-ended genre captioning
- Only two genre tags were included (Dubstep Growls and Chiptune waveforms), primarily to reinforce waveform behaviors
- Prompt quality matters β structured layered prompts outperform vague natural language
- Some timbre tags exert stronger influence than others
- Certain tag combinations may require iteration to achieve the exact musical role or timbral blend desired
- Percussion and drum sounds are outside the scope of this release
The model is also optimized around specific timing relationships between Bars, BPM, and generation duration.
For example:
- an 8-bar loop at 100 BPM β 19 seconds
If the generation duration is shorter than the musical structure implied by the prompt (for example requesting an 8-bar loop but generating only 5 seconds), the model may produce less coherent musical phrases.
The RC Stable Audio Fork automatically handles this timing alignment, making this workflow much easier.
Keybed-Specific Limitations
The Keybed model is designed around practical instrument ranges.
Some prompts simply stop making semantic or perceptual sense at extreme registers. For example, a sub-bass at C7 is no longer functioning as a bass, even if the model preserves some aspects of the original waveform or timbral character.
The same applies to many synthetic and acoustic sounds. As pitch rises into extreme upper registers, complex waveforms often become perceptually simpler and can collapse toward high-pitched pure-tone-like behavior. At the opposite end, very low pitches can become increasingly difficult to render consistently, and small pitch or waveform errors become more obvious.
For this reason, RC Stable Audio Tools clamps generation and export ranges to practical regions rather than forcing every generated sound across the full MIDI note range.
Prompt choice also matters. A timbre that has a natural low, mid, or high-register identity may become ambiguous when pushed several octaves outside that range. Some apparent "timbre drift" is therefore not only a model limitation; it is also a consequence of asking for a sound whose defining characteristics change as pitch moves far beyond where that sound is normally perceived.
In my internal testing, roughly 90% of generated instruments maintain a consistent relationship between pitch and timbral identity across the supported keybed ranges. This is a qualitative estimate rather than a programmatic benchmark, because there is no reliable automated metric for determining whether two notes share the same perceived timbre.
The remaining edge cases are most likely to appear with:
- extreme low or high registers
- unconventional bass prompts outside bass ranges
- strongly formant-dependent sounds
- heavily processed or hybrid timbres
- prompts whose intended instrument identity becomes ambiguous at the requested pitch
The default export ranges were chosen as a practical compromise between keyboard coverage, pitch accuracy, and timbral consistency.
License
This model is licensed under the Stability AI Community License. It is available for non-commercial use or limited commercial use by entities with annual revenues below USD $1M. For revenues exceeding USD $1M, please refer to the repository license file for full terms.
Companion Videos
Foundation-1 Overview
The original Foundation-1 video covering the model's design philosophy, structured prompting system, timbral control, and loop-generation workflow.
π₯ Watch the Foundation-1 Overview
Foundation-1.2 Samples & Keybeds Update
A companion video covering the new Samples and Keybeds checkpoints, the updated training strategy and general journey to getting this made can be found here.
π₯ Watch the Foundation-1.2 Update
Keybed Guided Demo
A hands-on walkthrough showing the Keybed model in use, cross-pitch timbral behavior, and the new three-layer instrument exporter.
π₯ Watch the Guided Keybed Demo
Final Notes
Foundation-1 is intended as a producer-facing model family for structured sample and instrument generation, designed to augment music production.
Its goal is to let users explore sound in new ways while retaining precise control over:
- what the sound is
- how it behaves musically
- how it changes across pitch
- how it sits tonally
- how it feels sonically
- how it fits into a production workflow
The original loop model focuses on structured musical material. The specialized Foundation-1.2 Samples and Keybeds checkpoints extend that same conditioning system toward sound design and playable instrument creation.
That combination of musical structure, instrument identity, timbral control, loop fidelity, and cross-pitch instrument generation is what defines the Foundation-1 family.
Model tree for RoyalCities/Foundation-1
Base model
stabilityai/stable-audio-open-1.0