Audiogram
A podcast-style audiogram: a static cover image, an animated waveform synced to the audio, episode title text, and burned-in subtitles for accessibility. Vertical 9:16 — the standard share format for Reels, Stories, and Shorts.
Complete JSON
{
"comment": "Podcast audiogram with cover, waveform, title, subtitles",
"resolution": "instagram-story",
"quality": "high",
"scenes": [
{
"elements": [
{
"type": "image",
"src": "https://cdn.json2video.com/assets/samples/podcast-cover.jpg",
"resize": "cover",
"position": "center-center"
},
{
"type": "audio",
"src": "https://cdn.json2video.com/assets/samples/podcast-clip.mp3"
},
{
"type": "text",
"text": "Episode 42\nBuilding in public",
"y": 150,
"style": "002",
"settings": {
"font-size": "75px",
"color": "white",
"font-weight": "800",
"text-align": "center",
"line-height": "85px"
}
},
{
"type": "audiogram",
"color": "#FF6B00",
"amplitude": 6,
"y": 1400,
"height": 200,
"width": 1000,
"x": 40
},
{
"type": "subtitles",
"settings": {
"style": "classic",
"position": "bottom-center",
"max-words-per-line": 4,
"font-size": 50,
"all-caps": false
}
}
]
}
]
}
How it works
The scene has no explicit duration — it inherits the length of the longest element, which is the audio clip. Whatever your podcast-clip.mp3 is, the video matches it exactly.
The image element fills the canvas as the visual backdrop. resize: "cover" scales the cover art until it fills the 1080×1920 frame (the instagram-story preset), and position: "center-center" centres it. A square cover (1080×1080) is scaled up to 1920×1920 and cropped on the left and right. To show square cover art uncropped instead, drop resize and position and place it explicitly: width: 1080, height: 1080, y: 420.
The audiogram element renders an animated waveform synced to whatever audio is playing in the scene — it doesn't own the audio, it visualises it. Properties:
color— bar colour. Match your brand accent.amplitude— how reactive the bars are to audio loudness. Start at 5-6; increase for quiet recordings, decrease for loud / compressed ones.height,width,x,y— position and dimensions in pixels.
The subtitles element generates timed text from the audio via JSON2Video's automatic transcription. style: "classic", max-words-per-line: 4, and position: "bottom-center" produce the standard reel-friendly caption look. The model defaults are good — whisper runs on the audio source automatically.
Cost note
The subtitles transcription is included at 0 credits today (see Credit consumption), so this audiogram costs only the rendering: 1 credit per second of output. Re-rendering the same audio during development still re-renders the video, so each render is billed for its duration.