Sources

Markdown

A source is one publication. A podcast, a newsletter, or a blog. Each source has a name, a type, a homepage, and the posts it has published.

A post is one item from a source. One newsletter issue, one podcast episode, one blog post. See Posts.

How we curate

Which sources do we add, and how do we decide?

We hold tens of thousands of publications. For each one we keep every post it has published since we started watching it, plus whatever archive it made available when we added it.

We add publications by hand rather than crawling. We work outward from publications we already index, looking at who they link to and who appears as a guest on them.

We leave out wire copy, aggregators and press-release syndication. That content is the same everywhere and has no author behind it. We also remove publications that turn out to be reposted wire copy or SEO filler after we added them.

If a publication you care about is missing, tell us and we will add it. Adding one is a small job.

Podcasts get transcribed

Where a show publishes its own transcript, we use it. Otherwise we transcribe the audio ourselves and identify the speakers. See Transcripts.

What you get on a source

Field What it holds
slug The key. Pass this back, not the name.
sputnik_url The publication's page on Sputnik.
name The publication's name.
type podcast, newsletter or blog.
description The publication's own description of itself.
homepage_url Its website.
avatar_url Its cover image or logo.
authority_score How much weight it carries in its space. Our assessment. Use it to rank a coverage report.
liveness Whether it is still publishing. Our assessment. Values below.
content_depth How substantial its posts are. Our assessment. A link roundup and a 4,000-word essay are different things to be mentioned in.
apple_id, spotify_id Its ids on those platforms, where it has them.
posts_count How many posts we hold.
feeds_count How many feeds sit behind it. Usually one.
first_post_published_at, last_post_published_at The range we cover.
language Two-letter language code, where the feed declares one.
explicit Whether the publication marks itself explicit.
genres Apple's genre list for it, where it has one.
owner_name, owner_email, author_name What the publication says about itself in its feed. Not verified.
twitter_handle Its X/Twitter handle, from the Substack leaderboard.
corpus_metrics What we've computed from its own posts: posting_cadence_per_week, posting_consistency (0-1, higher is more regular), feed_hosts, has_custom_domain, and computed_at. Recomputed weekly.

Asking for one source adds three more fields:

Field What it holds
topics What it has been covering lately. See Topics.
sponsors Brands heard on its recent posts. See Sponsors.
audience Audience/rank numbers, each with its provenance and history: subscriber counts, leaderboard rank. See Audience.

Both topics and sponsors come from the publication's recent posts, so both are empty for a publication we have only just started watching.

Its posts

A source's posts do not come back with the source. Ask for them separately, or search inside the publication.

/api/v1/posts?source=<slug>
/api/v1/posts/search?q=pricing&source=<slug>

See Posts for what you get on each one.

Our three judgment fields

authority_score, liveness and content_depth are our assessment rather than anything the publication told us. They are what keeps a corpus this size usable. Without them every list is dominated by whoever publishes most often.

Liveness values

  • active — publishing on schedule.
  • dormant — quiet lately. Many good newsletters go quiet for a month.
  • dead — stopped publishing. Hidden from results by default.
  • broken — its feed stopped working. We retry these.
  • pending — recently added, not yet assessed.

Type and platform

These are two different fields.

  • source_type is what the publication is: podcast, newsletter or blog.
  • publishing_platform is where it is hosted: substack, beehiiv, rss, bluesky or email.

A podcast hosted on Substack has source_type=podcast and publishing_platform=substack. Filtering on the platform when you meant the format is a common mistake.

What is excluded from results

Archived publications never appear. When we remove a publication for quality, its posts stop showing up in search and browse. There is no flag that brings them back.

Dead publications are hidden by default. Pass include_dead_sources=true to include them. This is useful when you are researching history rather than watching for news.

Freshness

We fetch continuously. An alert on a newsletter usually reaches you within minutes of it going out. Podcasts take longer, because transcription happens after the episode is published.

Every source carries last_post_published_at. Together with liveness, that tells you whether a publication is still going.