TRENDSCOUT

It reads your Twitter feed and 700 research papers every morning — and hands you the five worth building

Pipeline ArchitectPython / Local LLM / arXiv + Web2025 — NOW

The Problem

A one-person operation can't read 700 research papers a day. But the five that matter compound — the paper that ships as a feature next month, the technique that cuts inference cost in half, the niche that's about to be crowded. Miss them for a week and you're behind. Read them all and you're doing nothing else.

Attention is the scarcest resource in a small autonomous operation. Spending it on triage — deciding what to read before you can read anything — is the worst possible use of it. The studio needed a system that reads the firehose so the operator doesn't have to, surfaces only what's actionable, and is honest about the days when it couldn't do its job.

The Daily Drop

Trendscout is a four-stage pipeline that runs on a fixed schedule, lands a ranked brief before the operator wakes up, and checks its own delivery afterward. Two lanes run in parallel: an arXiv lane scanning six computer-science feeds, and a market lane that merges the operator's own X/Twitter “For You” feed with Reddit, Hacker News, GitHub trending, and ProductHunt into one cross-source stream, briefed per business unit — an AI-products studio, a generative-music brand, and the personal ideas backlog, each with its own focus areas baked into the prompt.

The Twitter lane is deliberately minimal: a read-only headless-browser pass over the For You timeline — 120 tweets captured, engagement metrics parsed, the top dozen threads expanded, and zero clicks, likes, or writes. The algorithm already spends all day learning what the operator cares about; Trendscout treats that curation as a free signal source instead of rebuilding it.

700 papers in. 5 build-worthy out. Ranked by a local model for $0/day — and honest about every deferred run.

The daily sequence:

Early morning — the arXiv pipeline runs, before the morning digest lands. Six feeds fetched, deduplicated, capped at 200 fresh items, scored by the local model.

Later — a summarizer turns the latest drop into a 5-bullet “why this matters” brief.

Midday — the market-trends digest lands on the operator's phone: niche signals, product trends, creator-economy movement per business unit.

After delivery — a health check verifies the digest actually sent. Delivery is measured, not assumed.

Weekly: a calibration job grades past niche predictions against real sales outcomes.

Pipeline Mechanics

L1 — Fetch & Dedup. Each run fetches roughly 688–742 raw items across six arXiv CS feeds. After deduplication: ~504–551 unique. Capped at 200 fresh. A real run header looks like this:

L1 fetched 688 raw → 504 unique → 200 fresh (feed_failures: 0/6)

The dedup state tracks 6,698 paper IDs with first-seen, last-seen, and version fields. Suppression window: 30 days. Version-aware: a v2 revision of a suppressed paper legitimately re-surfaces.

L2 — Local-Model Rank. A locally-hosted open-weights model scores every paper 0–10 on four axes: stack relevance, novelty, tractability (“shippable in under 1,000 lines in one session?”), and whether the paper extends one of the operator's last 30 active ideas. Zero cloud API calls. Zero marginal cost per paper. Processed in batches of 25, 180-second timeout, 2 retries. The inference rig powering this pipeline is the same one described in The Inference Rig.

L3 — BUILD-WORTHY Queue. Papers that clear all four thresholds — stack score ≥9, tractability ≥8, novelty ≥6, AND extending a live idea — get flagged BUILD-WORTHY and appended to a persistent queue. Hard cap: 5 per day, so a generous scoring run can't collapse the queue into noise. 40 entries accumulated to date, each with a one-sentence “why” and the idea numbers it extends.

Core pipeline: 1,271 lines across 5 modules. 15 dated runs spanning June 10 → July 14; 8 fully ranked, 7 honestly deferred (see below).

Craft Details

Craft 1 — The JSON salvage parser.

Local models truncate and malform JSON under load. A standard array-level parse fails the entire batch when one object is corrupt — which, at 25 papers per batch under a constrained context, happens constantly.

The ranker doesn't parse the model's output as a single array. It uses a balanced-brace scanner: it walks the raw text byte-by-byte, identifies object boundaries by tracking brace depth, and attempts to parse each object independently. A broken object around line 40 doesn't discard objects 1–39 and 41–25. Every intact object around the break is recovered.

Result: papers successfully scored per run went from 62% to 94%. The same broken model output, different parser, 32 percentage points of recovered signal.

A broken object around line 40 doesn't discard lines 1–39. The scanner recovers every intact object around the break.

Craft 2 — Fail-open never silently.

Every failure mode has a defined honest behavior. None of them silently serve stale or missing data as if it were fresh.

Some feeds down → proceed on partial data, exit code says so. ALL feeds down → abort entirely. A ranker with no data has nothing to rank; a partial run on zero items is not a run.

Local model unreachable → the drop doesn't silently skip. It writes DEFERRED — model unavailable into the digest. Seven of 15 runs deferred this way — the receipts show it plainly. The pipeline doesn't pretend it ran when it didn't.

Digest older than 48 hours → the morning report flags stale digest instead of serving old news as fresh. Every run header publishes the raw → unique → fresh counts and the feed failure count so the reader sees exactly what the pipeline saw.

House rule: a detector that inspected zero items reports NO-DATA, never PASS.

Seven of fifteen runs were deferred. The receipts say so. The digest doesn't pretend otherwise.

Craft 3 — Calibration against reality.

A research scout that surfaces “high potential” niches and never checks whether it was right is a confidence machine, not an intelligence one. Every week a calibration job grades past niche predictions against actual sales outcomes from each business unit.

Current status, honestly: INSUFFICIENT_DATA (n=1 graded). One prediction has enough follow-through data to grade. The calibration loop is live; the data to feed it is still accumulating. The gate exists before the data does — by design. An accuracy claim here would be a fabrication; the honest answer is the gate is running and waiting.

Craft 4 — The Plan button: from trend to plan in one tap.

A digest that only informs still leaves the operator with the expensive part: figuring out what a trend means for their stack. So every card in the drop carries a tappable 🧠 Plan this → action. Tapping it sends the card's ID back through the operator console, where a router looks the card up and has the local model draft a full implementation plan — mapped onto the operator's actual infrastructure, not generic advice. Plans are cached on disk: tap the same card twice, get the same plan instantly.

A real one: the feed surfaced an inference-runtime update claiming a 78% throughput gain for local models. One tap produced a plan that mapped the change onto the studio's own GPU fleet, flagged that the demo model differed from the models actually deployed, tiered the risk (touches live inference — benchmark side-by-side, never rip out what works), and split the work into an additive investigate phase and a separate cut-over decision. The operator approved it; it became a numbered item in the ideas backlog with a benchmark phase running. Dozens of plans have been drafted this way — each one a bridge from “interesting” to “actionable on our hardware.”

Read the feed → tap once → get an implementation plan for your own infrastructure. The distance from signal to action is one thumb.

Stack

PythonarXiv APILocal Open-Weights ModelJSON Salvage ParserShell / cronsystemdWeb Search APIHeadless Browser (read-only)Operator console actions

Result

What's real and counted:

688–742 raw items fetched per day across 6 arXiv CS feeds. 6,698 paper IDs in the dedup state, 30-day suppression window. 62% → 94% score rate after the JSON salvage fix. 40 BUILD-WORTHY entries in the queue to date, each with a one-sentence rationale and linked idea numbers. 1,271 lines of pipeline code across 5 modules. 15 dated runs, 8 fully ranked, 7 honestly deferred. $0 marginal cost per run — the local model runs on-premises hardware.

What's UNMEASURED (labeled honestly):

How often BUILD-WORTHY papers become shipped features — hit rate is not yet tracked. Prediction accuracy — calibration has n=1 graded, insufficient for a claim. Reader engagement with the daily drop — not instrumented by design (the brief arrives, it doesn't report back).

The attention cost of reading the firehose is zero. The intelligence cost of missing what matters is compounding. The pipeline takes the first problem seriously enough to be honest about the second.