Transcripts
MarkdownPodcast episodes are transcribed and split into speaker turns. A turn is one speaker's continuous stretch of talking, as a paragraph.
Where a show publishes its own transcript we use that. Otherwise we transcribe the audio ourselves.
Reading one
There are two shapes, picked with format. Both page by turn, so limit, cursor and total mean the same thing either way.
Markdown
The episode as readable text, laid out the way the Sputnik web page shows it. Use this when you want to read the episode, which is most of the time. Over MCP this is the default.
curl -H "Authorization: Bearer $SPUTNIK_TOKEN" \
"https://sputnikintelligence.com/api/v1/posts/<slug>/transcript?format=markdown"
{
"markdown": "**Ben Gilbert**\n\nSo the thing everyone gets wrong...\n\n**Jane Smith**\n\nRight, and we saw...",
"total": 412,
"next_cursor": "300",
"transcript_status": "fetched"
}
The Markdown itself reads like this:
**Ben Gilbert**
So the thing everyone gets wrong about this...
**Jane Smith**
Right, and we saw exactly that when we...
**Ad read** (ad)
This episode is brought to you by...
Turns
The same content as JSON, one object per turn. Use this when you need the structure: skipping the ad reads, counting who spoke most, or copying one turn exactly.
curl -H "Authorization: Bearer $SPUTNIK_TOKEN" \
"https://sputnikintelligence.com/api/v1/posts/<slug>/transcript?format=turns"
{
"items": [
{ "speaker": "Ben Gilbert", "kind": "speaker", "text": "So the thing everyone gets wrong..." },
{ "speaker": "Jane Smith", "kind": "speaker", "text": "Right, and we saw exactly that when..." },
{ "speaker": "Ad read", "kind": "ad", "text": "This episode is brought to you by..." }
],
"total": 412,
"next_cursor": "300",
"transcript_status": "fetched"
}
Each item is one turn.
| Field | What it holds |
|---|---|
speaker |
The person's name where we identified them. Otherwise the label the transcript used, such as Speaker 1. Null for narration on a single-voice show. |
kind |
What sort of voice this is. See below. |
text |
What was said, as one paragraph. |
Consecutive turns by the same speaker are merged, so you get the shape of the conversation.
This is the only way to read a transcript. The text is never included on the post itself. A long episode runs to hundreds of kilobytes.
How we store it is a plain text file with a speaker tag on each line. Neither shape above hands you that file. Both are rendered when you read, which is what lets us substitute names that were corrected after the episode was transcribed.
Speaker names
Where we worked out who was talking, speaker is their real name. Where we could not, you get the generic label the transcript used.
Names are substituted when you read, not when we store. A later correction or a people merge shows up straight away.
Kinds of turn
| Kind | Meaning |
|---|---|
speaker |
A named voice. |
unknown |
A voice we could not identify. speaker holds the generic label. |
ad |
A sponsor read. Skip these if you are looking for what the hosts said. |
announcer |
An intro or outro voice. Only on older transcripts. |
clip |
Audio played into the episode from somewhere else. Only on older transcripts. |
null |
Text with no label attached. |
A turn is marked ad when it reads out one of the post's sponsors the way an ad does ("brought to you by", a promo code, a URL). The speaker name stays on the turn: a host reading an ad is still the host. We aim for precision, so some sponsor reads will still arrive as speaker; treat ad as a reliable skip signal, not a complete one.
kind does not tell you who is the host and who is the guest. That is a property of the person on the post, not of the turn, and you get it from the post's people. We resolve it when you read, so a later correction is picked up straight away.
Not every episode has one
Filter with has_transcript:
/api/v1/posts/search?q=pricing&source_type=podcast&has_transcript=true
transcript_status comes back with the transcript too. If you get zero turns, it tells you whether the episode is not transcribed yet or transcribed and empty.
Paging
Turns are paged, 300 at a time by default. Follow next_cursor until it is null. total is the number of turns in the whole episode, so you know up front how much there is.
Accuracy
Transcription is machine-generated. Unusual proper nouns, company names and crosstalk are where it slips. If you are going to quote something publicly, listen to the audio at that point first. The audio file is on the post as enclosure_url.