Posts
MarkdownA post is one item from a publication. One newsletter issue, one podcast episode, one blog post. Posts are what you search, and what an alert matches against.
Every post belongs to one source.
Getting posts
# Search across everything
/api/v1/posts/search?q=%22Acme%22&window=month
# Newest first, for one publication
/api/v1/posts?source=<slug>
# One post in full
/api/v1/posts/<slug>
Search ranks by relevance and adds a snippet showing the match. Browse gives you a chronological feed. See Search.
Fields on every post
| Field | What it holds |
|---|---|
slug |
The key. Pass this back, not the title. |
sputnik_url |
The post's page on Sputnik. Link to this. |
title |
The post title. |
link |
The publisher's own URL. Different from sputnik_url. |
author |
As the publication declared it, unparsed. For people we identified, use people. |
published_at |
When the publisher dated it. Null when the feed carries no date. |
source |
The publication it came from: slug and name. |
excerpt |
A short summary, where the feed gives one. |
snippet |
Search results only. The matched text, with <b> around the hit. |
categories |
The publication's own tags. Sparse and unnormalised. |
transcript_status |
Podcast episodes. Whether a transcript exists yet. |
Fields you only get on one post
These are too big or too expensive to put on every row of a list.
| Field | What it holds |
|---|---|
body |
What the publisher wrote, as markdown with links kept. For a podcast episode, the show notes. |
enclosure_url |
The audio file, for a podcast episode. |
people |
Who we credited on it, each with their role. See People. |
engagement |
Reactions, comments and restacks from the publishing platform, where we have them. Null when we don't. |
original_metadata |
Extra values from the feed, such as duration_seconds. The feed's own episode summary is left out: it repeats body. |
Engagement
engagement holds the counts a publishing platform reports on a post, where we have them today: only Substack. It is null for every other platform, and for a Substack post we haven't fetched from the archive yet.
| Field | What it holds |
|---|---|
source |
Where the numbers came from. |
fetched_at |
When we last read them. |
reaction_count |
Total reactions (likes, etc). |
reactions |
The breakdown by reaction, where the platform gives one. |
comment_count |
Top-level comments. |
child_comment_count |
Replies to comments. |
restack_count |
Times readers reshared it. |
post_type |
The platform's own type for it, e.g. newsletter or podcast. |
paywall_status |
Who could read it: e.g. everyone or only_paid. |
is_audio |
Whether it carries audio. |
cover_image_url |
The cover image the platform shows for it. |
word_count |
Length, in words. |
section_name |
The publication's own section, where it has sections. |
Only the latest numbers are kept — no history.
Reading the content
body holds what the publisher wrote, as markdown with links kept. For a newsletter or blog post that is the article. For a podcast episode it is the show notes: the description, guest links and sponsor codes. It comes back when you ask for one post. Where we have no markdown for a post, body is plain text.
Podcast episodes also have a transcript: what was said on the audio. It is a separate request because an episode runs to hundreds of kilobytes. See Transcripts.
/api/v1/posts/<slug>/transcript
Two posts that look the same
A publication sometimes ships the same piece more than once, for example a corrected version, or the same episode on two feeds. We group those and return one post. The one you get is the version we consider canonical.
Dates
published_at is when the publisher says it went out. It is not when we found it. Some feeds carry no date at all, and those posts have published_at: null.
Alert hits carry their own created_at, which is when we found the match. See Alerts.