Lantern: an agent that keeps watch over your repository
Most coding agents wait for you to type something. Lantern doesn't. It is a local daemon that wakes up on a schedule, works through a standing checklist against your repository (scan the logs, read the bug inbox, run the tests), and files what it finds as tickets in a ledger it owns. It hands the tickets it can handle to Claude Code sessions and escalates to you only what needs a decision.
You check in when it suits you. One terminal command, or a local web page, shows what it noticed, what it did, what it spent and what is waiting on you.
This post covers how it works and why it is built the way it is. The short version: an agent that runs while you're away has to be cheap when nothing is happening, honest about what it didn't check, and unable to do real damage even when its model is confused or being manipulated.
The pulse
Everything Lantern does happens inside a pulse, a heartbeat that fires every few minutes or every hour, depending on the agent:
heartbeat --> gates ---> collect ---> Claude turn ---> reap children
(budget, (logs, (triage, (record
quiet inbox, staff, results,
hours...) tests) escalate) push
branches)
- Gates decide whether the pulse runs at all. Lantern skips it if the agent is paused, it's inside configured quiet hours, the daily budget is spent, or the repository has uncommitted changes. That last one means it never builds on top of your half-finished work.
- Collectors gather evidence in plain TypeScript, without a model.
logs.scanreads plain-text or NDJSON log files,inbox.bugsreads exported bug-report emails, andci.watchruns the test suite on the default branch in a scratch worktree, parses the JUnit XML, and tells a failing test apart from a flaky one. - The orchestrator turn is one Claude Code session that reads a digest of the collectors' findings plus the open ledger, and decides what to do: triage, staff a ticket, or escalate a question.
- Reaping collects results from finished child sessions, records them on their tickets, and publishes any branches they produced.
Collectors run in their own phase, before the model is involved. That keeps their time and cost out of the model's budget, lets them be tested without a live model, and means the turn reasons about findings that are minutes old rather than kicking off its own scans.
A ledger, not a chat
An agent that runs every fifteen minutes can't hold its memory in a context window. Lantern's memory is a SQLite database: tickets, ticket events, pulses, sessions, escalations and spend. Each pulse starts fresh and reads the state of the world from the ledger.
The ledger is also the source of truth for the agent's actions. The model changes things by calling tools on Lantern's own MCP server, such as staff_ticket, escalate and pulse_finding, and each call is written to the database in a transaction. The model's final text is only a summary. Lantern never parses prose to work out what happened.
A few ledger features keep the queue worth reading:
- Fingerprints and dedupe. Every error gets a fingerprint. A repeat of an open ticket adds to that ticket instead of opening a new one, and a fingerprint closed as a duplicate follows its links to the ticket that is still open.
- Noise suppression with an escape hatch. Dismissing a ticket as noise keeps that error from being filed again for 30 days, unless its rate jumps tenfold. A quiet error that suddenly gets loud is news again.
- Prior art. When a new ticket arrives, full-text search over past tickets, ranked by matching fingerprints and code locations, shows the session what was already tried.
- Triage first. Nothing a collector files gets worked on until a human accepts it. The agent can notice anything. It can only act on what you've agreed is real.
Three tiers of trust
Accepted tickets go to Claude Code child sessions, each started at one of three capability tiers:
read_onlyinvestigates and reports. It can read code, search, and run read-only git commands, and nothing else.write_branchworks in its own git worktree on a fresh branch. It can edit files, run the agent's configuredtest_commands, and commit. Lantern then pushes the branch and, on GitHub, opens a draft pull request for you to review.releaseis likewrite_branchplus a tag, and it never starts without your explicit approval.
The tiers are enforced by the session's tool set, not by instructions in the prompt. Claude Code's --allowedTools flag only pre-approves tools. It doesn't take the others away. So each tier also starts the session with a restricted list of built-in tools and an explicit deny list. If a tool isn't on a session's list, the session can't call it, whatever the prompt says.
The PetClinic example shows the escalation in practice. A bug report says looking up a missing owner returns HTTP 500 instead of 404. You accept it, and a read_only child diagnoses it and writes its findings on the ticket. You comment "fix it", and on the next pulse a write_branch child applies the fix on its own branch, runs the project's tests, and the branch lands on the remote for review.
Guardrails below the prompt
Lantern's inputs include log lines, test output and emails from strangers, so it assumes any of them could be an attack. Its safety rules live in code, below the model:
- Evidence, never instructions. Collected text reaches the model only inside a marked untrusted block. Each block's delimiters carry a random nonce, so an email that contains a fake "end of untrusted data" line can't close the block early. In the web UI the same content renders as plain text.
- No dangerous pushes. Child sessions can't push at all. Lantern does every push itself, through a pre-push guard that refuses
mainandmaster, force-pushes and moving an existing tag. - Hard spend ceilings. Every turn and every child has a cost ceiling, which defaults to $1.50 and is passed straight to the harness. Each agent also has a daily ceiling. An agent that hits it trips and stays stopped until a human runs
lantern resume. - A kill switch.
lantern pausestops all pulses and kills any running child sessions immediately. - Settings split by risk. You can change the heartbeat, concurrency, budgets and quiet hours from the web page. Anything that decides what a child may run or touch can only be changed in the agent's file on disk.
Paying for news, not for idling
An always-on agent that runs a model turn every pulse burns money doing nothing. On the PetClinic demo, an idle turn costs $0.10–0.16. At a 15-minute heartbeat that adds up to about $12 a day of "nothing new here".
So the collectors, gates and reaping run every pulse, but the expensive part, the Claude turn, runs only when something has changed since the last one. Lantern runs a turn when a collector filed a ticket or hit an error, a human acted, an escalation was answered, a child session finished, or you pressed "Pulse now". It also runs one every couple of hours regardless, so the agent still looks around now and then. An idle agent costs close to nothing.
Skipped is not clean
The most dangerous thing a monitoring agent can say is "all clear" when it didn't actually look. Lantern treats "clean" as a claim about coverage and won't make it without evidence:
- A log window with no data is reported as skipped, never clean.
- A pulse's status is worked out from what it actually did, in this order: escalated, then acted, then skipped, then clean. A pulse counts as clean only when at least one check actually ran and found nothing.
- Status comes from rows in the ledger, not from wording. An early version read status from the model's summary and marked a pulse "escalated" just because the summary said it couldn't escalate. Since then, the database is the only thing that counts.
Thirty days of these honest numbers feed lantern metrics: coverage, how often each collector's tickets turn out to be real, and cost per fix.
Checking in
You can run Lantern entirely from the terminal:
lantern status # agents, last 24 h of pulses, spend
lantern triage # accept or dismiss what the collectors filed
lantern inbox # questions and release approvals waiting on you
lantern sessions --live # running children; lantern kill <session>
lantern logs <pulse-id> # everything one pulse did
lantern metrics # coverage, precision, cost per fix
Or with lantern ui, which serves a local page on 127.0.0.1, protected by a token in the URL. It has an overview of every agent's health, a triage queue, ticket histories with their prior art, live-streamed session transcripts, an inbox for escalations, a settings panel, and a doctor view that checks your setup. It updates live over server-sent events, so you can leave it open and watch a pulse happen.
Each agent is described by a single Markdown file: YAML front matter for the configuration (repository path, heartbeat, budgets, log sources, test commands) followed by plain-language sections that become its instructions. Before an agent does anything for real, lantern pulse --now --dry-run checks the gates and prints the exact prompt it would get, without running a model or filing anything.
Try it
Lantern is MIT-licensed, has one runtime dependency (better-sqlite3), and runs its TypeScript directly on Node 24 with no build step. More than 500 tests cover it, from the ledger and the gates to adversarial log and inbox fixtures that try to break out of the untrusted block.
git clone https://github.com/cleiderg/lantern.git
cd lantern && npm install && npm link
lantern init && lantern doctor
The fastest way to see it work is the Spring PetClinic example. It runs a real Java app that breaks on demand, gives it a small inbox of bug reports (one of them isn't a bug), and points Lantern at a local git remote so nothing can reach GitHub. Start the error traffic, open the UI, and watch tickets appear, get triaged, and turn into reviewed branches.
Source and docs are on GitHub.