Sources
MarkdownA source is one publication. A podcast, a newsletter, or a blog. Each source has a name, a type, a homepage, and the posts it has published.
A post is one item from a source. One newsletter issue, one podcast episode, one blog post. See Posts.
How we curate
Which sources do we add, and how do we decide?
We hold tens of thousands of publications. For each one we keep every post it has published since we started watching it, plus whatever archive it made available when we added it.
We add publications by hand rather than crawling. We work outward from publications we already index, looking at who they link to and who appears as a guest on them.
We leave out wire copy, aggregators and press-release syndication. That content is the same everywhere and has no author behind it. We also remove publications that turn out to be reposted wire copy or SEO filler after we added them.
If a publication you care about is missing, tell us and we will add it. Adding one is a small job.
Podcasts get transcribed
Where a show publishes its own transcript, we use it. Otherwise we transcribe the audio ourselves and identify the speakers. See Transcripts.
What you get on a source
| Field | What it holds |
|---|---|
slug |
The key. Pass this back, not the name. |
sputnik_url |
The publication's page on Sputnik. |
name |
The publication's name. |
type |
podcast, newsletter or blog. |
description |
The publication's own description of itself. |
homepage_url |
Its website. |
avatar_url |
Its cover image or logo. |
authority_score |
How much weight it carries in its space. Our assessment. Use it to rank a coverage report. |
liveness |
Whether it is still publishing. Our assessment. Values below. |
content_depth |
How substantial its posts are. Our assessment. A link roundup and a 4,000-word essay are different things to be mentioned in. |
apple_id, spotify_id |
Its ids on those platforms, where it has them. |
posts_count |
How many posts we hold. |
feeds_count |
How many feeds sit behind it. Usually one. |
first_post_published_at, last_post_published_at |
The range we cover. |
language |
Two-letter language code, where the feed declares one. |
explicit |
Whether the publication marks itself explicit. |
genres |
Apple's genre list for it, where it has one. |
owner_name, owner_email, author_name |
What the publication says about itself in its feed. Not verified. |
twitter_handle |
Its X/Twitter handle, from the Substack leaderboard. |
corpus_metrics |
What we've computed from its own posts: posting_cadence_per_week, posting_consistency (0-1, higher is more regular), feed_hosts, has_custom_domain, and computed_at. Recomputed weekly. |
Asking for one source adds three more fields:
| Field | What it holds |
|---|---|
topics |
What it has been covering lately. See Topics. |
sponsors |
Brands heard on its recent posts. See Sponsors. |
audience |
Audience/rank numbers, each with its provenance and history: subscriber counts, leaderboard rank. See Audience. |
Both topics and sponsors come from the publication's recent posts, so both are empty for a publication we have only just started watching.
Its posts
A source's posts do not come back with the source. Ask for them separately, or search inside the publication.
/api/v1/posts?source=<slug>
/api/v1/posts/search?q=pricing&source=<slug>
See Posts for what you get on each one.
Our three judgment fields
authority_score, liveness and content_depth are our assessment rather than anything the publication told us. They are what keeps a corpus this size usable. Without them every list is dominated by whoever publishes most often.
Liveness values
- active — publishing on schedule.
- dormant — quiet lately. Many good newsletters go quiet for a month.
- dead — stopped publishing. Hidden from results by default.
- broken — its feed stopped working. We retry these.
- pending — recently added, not yet assessed.
Type and platform
These are two different fields.
source_typeis what the publication is:podcast,newsletterorblog.publishing_platformis where it is hosted:substack,beehiiv,rss,blueskyoremail.
A podcast hosted on Substack has source_type=podcast and publishing_platform=substack. Filtering on the platform when you meant the format is a common mistake.
What is excluded from results
Archived publications never appear. When we remove a publication for quality, its posts stop showing up in search and browse. There is no flag that brings them back.
Dead publications are hidden by default. Pass include_dead_sources=true to include them. This is useful when you are researching history rather than watching for news.
Freshness
We fetch continuously. An alert on a newsletter usually reaches you within minutes of it going out. Podcasts take longer, because transcription happens after the episode is published.
Every source carries last_post_published_at. Together with liveness, that tells you whether a publication is still going.