Quickwit on Laravel Cloud with Rust, twelve days in
Peter Van Dijck · October 7, 2026 · 16 min read
Follow us on XTwelve days ago I moved our search engine, Quickwit, from a DigitalOcean droplet to Laravel Cloud's Rust beta, and wrote a very excited post about it. Short version: the Rust part is boring in the best way, and having everything in one environment with one vendor is really awesome. What went wrong, went wrong between my Laravel app and the Rust box.
A word on Quickwit. It's built for append-only data, and lots of it: documents arrive, get indexed and never change. That's us. And because nothing changes, it can keep its whole index in an S3-style bucket and search it from a small, stateless box.
The other character in this story is Hone, an open source, self-hosted replacement for Nightwatch. It went live the same day as the move, and most of the numbers in this post come from it.
Let's go through it. Numbers, incidents, the stuff I got wrong, and three things I'd love the Laravel team to add.
The setup is a small ARM box
Quick recap for people who didn't read the first post. Sputnik indexes podcasts and newsletters, and search runs on Quickwit 0.8.2, a search engine written in Rust. Quickwit keeps its index in an S3-style bucket instead of on local disk, which is cheap and fast (because we have Very Large Amounts of data).
Here's what's running:
- One Laravel Cloud app called
quickwit-cloud, on the Rust runtime (beta, not in the public docs yet). - One instance,
flex.m-2vcpu-2gb, in us-east-2. ARM (Graviton,aarch64). Minimum 1 replica, maximum 1 replica. - All index data and the metastore on Cloudflare R2, in the app's Cloud bucket. Nothing durable lives on the box.
- A 144-line Rust program in front of Quickwit that handles auth, the health check and keeping Quickwit alive. More on that below.
- On the Laravel side,
php artisan search:ingest-outboxruns as a background process on the main app. It reads finished search documents from the bucket and posts them to Quickwit.
What it replaced: a DigitalOcean s-2vcpu-4gb droplet at $24/month, managed through Forge, with nginx, two daemons and a firewall.
The index today is 2,597,955 documents in 39 splits. That's 11.7 GB of split files on R2, built from 18.3 GB of uncompressed text. Two big splits hold about 0.9 million docs each, and then there's a tail of tiny ones (1 to 6 docs) from the minute-by-minute ingest, waiting to be merged.
2 GB of RAM is plenty to search a 12 GB index
I started on 4 vCPU / 4 GB because that's roughly what the droplet had. That same evening I dropped it to 2 vCPU / 2 GB. Search times didn't change.
Quickwit's memory on the 2 GB box sits at 90 to 140 MB RSS. That includes right after running an aggregation over the whole index. On the droplet it used about 257 MB, so the 4 GB I was paying for there was sized for the caches I had configured, not for anything Quickwit was really using.
This works because Quickwit never loads the index. It reads split footers and fast fields from R2 as it needs them and keeps them in caches. Mine are set in the wrapper: 400 MB for fast fields, 200 MB for split footers, 64 MB for partial requests.
So memory is not the constraint. The constraint is how far away R2 is, and how often you're reading a split for the first time.
The first read of a split costs one to three seconds, and after that it's fast
I ran these today from Madrid, so the network adds about 170 to 250 ms on top. The numbers below are Quickwit's own time, which the wrapper reports in an X-Quickwit-Ms header.
| Query | First run (cold) | Repeat runs |
|---|---|---|
| Count, all-time ("openai") | 661 ms | 4–5 ms |
| Weekly histogram, 1 year | 146 ms | 4–5 ms |
| Term, last 7 days, 20 hits + snippets | 1,485 ms | 284–315 ms |
| Term, all-time, 20 hits + snippets | 2,767 ms | 398–515 ms |
| Phrase, all-time, 20 hits | 1,198 ms | 172–305 ms |
| Sorted by date, 50 hits | 1,181 ms | 139–501 ms |
| 100 hits + snippets, all-time | 939 ms | 530 ms |
| Terms aggregation on source, whole index | 730 ms | ~574 ms |
| Terms aggregation on people (size 1000), whole index | 1,469 ms | 1,234 ms |
What "cold" means here
Quickwit doesn't have one big index. It has splits. A split is a small, complete index in a single file that never changes once it's written. Every time the indexer commits a batch it writes a new split to the bucket, and in the background it merges small ones into bigger ones. That's why we have 39 of them: two huge ones and a tail of tiny ones.
Quickwit never downloads a whole split. When a query comes in, it first skips the splits that can't match, for example because their time range is outside the query. For each split that's left it fetches the footer, which is a map of where everything sits in the file. Then it asks R2 for only the byte ranges it needs: the entries for the search terms, the columns for sorting or aggregating, the stored text for the hits it returns.
Cold means none of that is in memory yet. One query becomes a handful of round trips to R2 per split, and R2 is about 118 ms away. Warm means the footers and columns are already in the caches.
The stored text isn't cached, which is why queries that return hits stay at a few hundred milliseconds when warm, while counts drop to 4.
Three things to read out of that.
Counts and histograms are basically free once warm. Four milliseconds to count every mention of a term across 2.6 million documents, on a box with 2 GB of RAM. I still find that a bit silly.
Cold is where you pay. The first touch of a split means fetching from R2, and that's one to three seconds. A restart empties the cache, so the first few queries after a restart are slow ones.
Whole-index term aggregations don't get warm. They scan fast fields across every split, every time, so they stay at 0.5 to 1.5 seconds. If you're building a "top sources for this term" feature, that's your floor.
Warm speed is the same as it was on the droplet. Region matters more than box size: R2 was about 118 ms away from any US region when I measured in August. I also tested Hetzner in Helsinki back then, where R2 was 259 ms away, and real queries came out about 2.5x slower. I haven't compared us-east-2 with DigitalOcean's NYC3 head to head.
The hidden cost is payload size. I store body and transcript in the index so Quickwit can return highlighted snippets, which means 20 hits is 259 to 414 KB and 100 hits is 1.1 MB. Cloudflare's edge gzips it (414 KB becomes 148 KB on the wire), but it's still a lot of bytes for a search result page.
One query failed outright with a 500: a per-day histogram over all time with a nested terms aggregation. My guess is Quickwit's bucket limit, since some posts have far-off dates. I haven't verified that, and I don't think it has anything to do with Cloud.
Autoscaling is off, and so is scale to zero
In the first post I got very loud about scale to zero. I have to walk that back for this app.
It also never sleeps. The ingester posts a batch about every 100 seconds, so the box is never idle long enough to hibernate. Hibernation is still switched on in the environment, which is wrong for a search node, and turning it off is on my list.
Deploys are the weak spot. A Cloud deploy briefly runs the old instance and the new one side by side. For a web app that's what you want. For Quickwit that's two writers. My rule is: deploy this app rarely, and never automatically. There have been four deploys, all on the day of the move.
The platform restarts things on its own schedule. On October 5 at 15:18 UTC the instance restarted without a deploy from me. Quickwit was serving again about 2.5 seconds after start, which is the upside of having nothing on the box. The downside of one replica is that search is down for those seconds.
The Laravel side does scale, from 1 to 4 web replicas. They take turns running the ingester through a cache lock, and you can see it in the Cloud logs: ingest requests arrive from three different IPv6 addresses.
If I ever need more search capacity, the path is read-only searcher nodes (--service searcher --service metastore) as a second app, with autoscaling on, while the one writer stays pinned. I ran that as a read-only pilot before the cutover and it works.
There's a 20 second limit on the gateway
Cloud's gateway gives up on a request after about 20 seconds and returns a 504. Your app doesn't get told. It carries on, finishes the work, and logs a 200.
I hit this on day one. The ingester was posting with commit=wait_for, which makes Quickwit hold the request open until the documents are searchable. That took longer than 20 seconds, so the gateway answered 504, Laravel saw a failure and retried, and Quickwit happily ingested the same batch again. About 30 minutes of posts are in the index three times. Search dedupes by post_id, so nobody sees it, but it's in there.
The fix was commit=auto, plus a 70-second wait before the ingester moves its checkpoint forward.
That should have been the end of it. Then I looked in Hone at the Laravel app's calls to Quickwit over the 12 days: 1,350 sampled requests, average 401 ms, and p95 20,062 ms. The p99 and the max were within five milliseconds of that.
A p95 that lands on 20 seconds to the millisecond is a cutoff. I read it as "5% of our calls to Quickwit are dying at the gateway" and got worried.
It was one bad day
It wasn't 5%. Hone's p95 over a long window is the worst day, split per deploy, so one bad group sets the number for the whole 12 days. Splitting by deploy found it:
- All the 20-second calls sit in one Laravel deploy, which ran from October 4, 09:00 to October 5, 09:06 UTC.
- That deploy had 50 calls with a 1.95 s average, which works out to about 4 or 5 calls cut off at the gateway.
- They were probably MCP searches.
POST /mcppeaked at 20.5 s in that same deploy, so most likely somebody's agent waited 20 seconds and got empty results. - Every deploy since is clean. The slowest call in each is 2.4 to 3.2 s, and the last 24 hours average 216 ms. Six hours of today's production logs show no Quickwit failures.
- The ingester is healthy. The gaps between ingest calls come in steps of about 30 s, one per empty polling round, never the extra 20 s a timeout would add. No retries in 12 hours.
I can't tell you what caused it
The Quickwit app's Cloud logs start at its October 5 restart. Everything before that is gone, while the Laravel app's logs from October 4 are still there. So it looks like the Quickwit app's logs get wiped when the instance restarts, possibly because the Rust runtime is in beta. I haven't confirmed that with Laravel.
They wouldn't have shown the failures anyway. The gateway returns the error to the caller while Quickwit finishes the request and logs a 200. In 12 hours of its access log there are 294 ingests and 71 searches, every one a 200.
The likely causes are the brief R2 connection drops Quickwit logs a few times a day, or a slow first read of a big chunk of the index. Neither is proven.
My own code made it harder. QuickwitClient turns any failure into null, so a search cut off at 20 seconds looks exactly like a search with no hits. That's the first thing I'm fixing: log the path and duration on every failed call, and make the client fail loudly.
What I learned about Hone
- Long-window percentiles mislead. A percentile built from the worst day hides that the trouble sat in a single day. Split by deploy.
- The background ingester reports late. It runs as one command for about a day, so Hone only gets its calls when it exits at a deploy, about 500 at once, and only for the 10% of runs that are sampled.
- Hone loses some data itself. Today's production logs have 8 timeouts on posts to Hone, which has a 500 ms limit.
The Rust part took 144 lines
If you want to try this yourself, here's everything Rust-specific I know.
You can fake it with an empty go.mod. Before I had beta access, Cloud refused a plain Cargo repo with "unsupported framework". Adding an empty go.mod made Cloud treat it as a Go app. The build step installed rustup, compiled, and copied the binary to ./app. About 30 seconds, and it worked. I'm not recommending it, but it's nice to know the platform is that flexible underneath.
With the beta it's just a Cargo repo. Cloud accepted it and set the build command to cargo build --release on its own. Hello world deployed in 41 seconds.
It's ARM. Builds run on aarch64. I don't compile Quickwit, which is a big workspace. The build script checks uname -m and downloads the matching release binary (quickwit-v0.8.2-aarch64-unknown-linux-gnu) from GitHub. The only thing that gets compiled is my wrapper. A full deploy, push to main through to live, takes about 52 seconds.
Why there's a wrapper at all
Cloud wants one process listening on $PORT with a health endpoint. Quickwit's REST API has no auth. So something has to sit in front, and that something replaces what nginx and a Forge daemon did on the droplet.
It's one file, src/main.rs, 144 lines, two crates (tiny_http and ureq). It does four things:
- Writes Quickwit's node config into the temp dir at startup: node id, cluster id, R2 endpoint, cache sizes.
- Starts the Quickwit binary on
127.0.0.1:17280, and restarts it 5 seconds after any exit. A supervisor loop. - Serves
[::]:$PORT. Answers/upwith a 200 for Cloud's health check, and returns 404 to anything without the bearer token. 404 and not 401, so a scanner doesn't learn there's something here. - Forwards everything else to Quickwit and adds the
X-Quickwit-Mstiming header.
It's plain blocking code, one thread per request, no async runtime. At our traffic that's fine.
What the wrapper doesn't do
The shortcuts I took:
- It reads whole request and response bodies into memory. No streaming. With 1.1 MB responses that's OK; with 100 MB it wouldn't be.
- It forwards only
Content-Type. Cloudflare still gzips responses at the edge, so I haven't missed the rest. - Its upstream timeout is 120 seconds, which is pointless when the gateway cuts at 20.
- The token check is a plain string compare, not constant-time.
And Quickwit itself is the reason any of this fits on a small box. It's Rust all the way down, with tantivy doing the indexing: about 140 MB of RSS for 2.6 million documents and a 2.5-second start.
The smaller stuff that bit me
None of these are dealbreakers. All of them cost me time.
There's no persistent disk. Quickwit's local disk is wiped on every restart. Everything durable is on R2, so that's fine.
Private networking is a Private Cloud feature. Laravel does have internal networking between apps, but it's part of Private Cloud, which is custom-quoted and aimed at teams spending $2,000 or more a month. On a regular plan, the Laravel app talks to Quickwit by going out through Cloudflare's public edge and back in. The bearer token in the wrapper is the only protection. It also means those calls count as bandwidth. Our organization is at 95% of the period's bandwidth allowance with 10 days to go. I don't know how much of that is Quickwit, but search responses are big and ingest bodies are small, so I have a suspect.
Renaming the app changes its URL. And the new URL only worked after a redeploy. That was about 30 minutes of downtime on moving day.
R2 blips. Four times in 12 hours Quickwit logged "metastore service is unavailable … connection closed before message completed". It recovered by itself every time.
Logs are hard to read in bulk. cloud environment:logs returns at most 100 lines per query. Pulling 12 hours took 36 calls, and 9 of those windows were still cut off.
There's an nginx in front of the Rust app. Every ingest logs an nginx warning, "client request body is buffered to a temporary file". So even on the Rust runtime, your binary isn't the first thing a request touches.
A CLI gotcha. cloud environment:variables --action=set reports success but doesn't change a key that already exists. --action=append does.
Everything is in one place now.
After all that complaining, here's why I'm staying.
The Laravel app, the bucket, the telemetry and now the search engine all live in one Cloud organization. One vendor, one CLI, one dashboard, one bill. When something looks off I look in one place, and when I want to change something it's the same commands I already use for the Laravel app.
For a solo founder that's really awesome, and it matters more than any number in this post. Day to day it looks like this:
- Nothing to look after. No droplet, no Forge, no nginx config, no SSH, no OS patching.
- A stateless node. A restart or a rebuild costs a cold cache and nothing else. Ready in about 2.5 seconds.
- Ingest is steady. A POST about every 100 seconds, small batches of 3 to 5 documents, all 200s.
- Deploys are a
git push. Cloud runscargo build --release && sh build-quickwit.shand it's live in about 52 seconds. - Warm queries are as fast as before, on half the RAM.
It's also cheaper: $4.55 from September 25 to October 7, so about $11 a month against $24 for the droplet. That money is trivial and it's not why I'd do this.
I used to have a search server. Now I have a repo with one Rust file and a shell script in it.
Three things I'd ask Laravel for
The Rust runtime is a beta, and for a beta it's in very good shape. The gaps I ran into are all about running something stateful-ish and long-running next to a Laravel app, which is exactly what a search engine is.
- A deploy mode that stops the old instance before starting the new one. For anything with a single-writer rule, the overlap is the dangerous moment.
- Internal networking on the regular plans. Private Cloud has it: apps in the same cluster talk over a private network with an internal URL. I'd love a version of that for two apps in one ordinary organization, so app-to-app calls don't leave through the public edge and don't eat bandwidth.
- A gateway timeout you can raise per app. 20 seconds is a sensible default for web requests. It's tight for a search engine doing a commit.
Would I do it again? Yes. Laravel Cloud is just so nice, so easy to use, and most importantly, they are doing the Right Things for us users.
Onwards!
More posts