Documentation

What we index

Newsletters, podcasts and blogs — publications where a named person with an audience is writing or talking. Tens of thousands of them, and every post each one has published since we started watching it, plus whatever archive it made available when we added it.

What we are not is a news wire. Wire copy, aggregators and press-release syndication are the opposite of the signal here: they are the same text everywhere, written by nobody in particular. The value is in a named host saying what they actually think.

How something gets in

Mostly by us adding it deliberately. We work outward from publications we already index — who they link to, who guests on whose show — and we curate rather than crawl. A publication that turns out to be reposted wire copy, an SEO farm, or a brand blog with no author gets removed again.

If something you care about is missing, tell us. Adding one publication is a small job and we would rather do it than have you working around a hole.

Podcasts get transcribed

A podcast is audio, which is not searchable. Where a show publishes its own transcript we use that; otherwise we transcribe it ourselves, and identify who is speaking. That is what makes a two-hour conversation something you can search a brand name in.

Not every episode has one — see Transcripts for what to expect and how to filter.

How fresh it is

We fetch continuously, and an alert on a newsletter usually reaches you within minutes of it going out. Podcasts are slower, because transcription happens after the episode drops.

Every publication carries a liveness — whether it is still publishing — and last_post_published_at. A corpus this size always contains shows that quietly ended two years ago, and those two fields are how you tell.

What gets excluded from results

Archived publications never appear. When we remove something for quality, its posts stop showing up in search and browse. No flag brings them back.

Dead publications are hidden by default. A publication we have marked dead has stopped publishing. Its posts are excluded unless you pass include_dead_sources=true — useful when you are researching history rather than watching for news.

Things it is worth knowing

We hold full text where we can get it. Many newsletters put their whole issue in the feed. Paywalled ones give us the preview, and that is what we index.

Sponsor mentions are ads that ran, not deals a show signed. Podcast ads are often inserted programmatically by the platform, so a show's transcript can read out a brand its publisher never sold to. See Topics and sponsors.

We credit fewer people than exist. Feeds are inconsistent about authorship, so we only record a person when we are confident. See People.