AI NewsDev toolsAnnouncement

Leviathan indexes large logs so an agent gets a few cited records back instead of a million rows

Leviathan is a new Apache-licensed Rust tool that turns a log file, JSONL export or SQLite table into a full-text index so an AI agent gets a short list of cited rows for each question instead of streaming the whole dataset into its context.

AI News

Editorial2 min read

LinkedInX
GitHub social card for the Leviathan repository

Image: GitHub

Why it mattersAny team whose agent asks questions over logs, tickets or event history pays for every row it has to read, and a short cited list keeps the token bill and the latency fixed while the data grows.

An AI coding agent asked to find one incident in a year of logs still reads the whole year, because the shell tools it uses do not know what matters. Leviathan is a new open-source tool that puts a full-text index between the agent and the data so each question comes back as a short list of cited rows.

The developer elstongun published Leviathan on 5 October 2026 as a single static Rust binary under Apache 2.0. The repository had 380 stars one day after it was created. It converts JSONL, JSON, CSV or TSV, SQLite databases, or the output of any database command line into a single SQLite index on disk, which agents then query in plain words. A leviathan wrap command sets it up for Claude, Cursor, Codex, VSCode, Gemini or Windsurf.

The numbers it claims

The project's own benchmark, documented in docs/BENCHMARKS.md, runs 1,200 questions against a synthetic maintenance log of one million records and 678 MB. The median answer uses 436 tokens from Leviathan against 107,122 from the best grep strategy the author tested, a 245 times difference. The median latency is 33 ms against 92 ms. Across the 1,200 questions, the relevant record is in the top five results 99.0 percent of the time and ranked first 98.5 percent of the time. The worst question in the set used 602 tokens through Leviathan and 9.7 million through grep.

elstongun states on the repository that the dataset is one example and nothing in the tool is specific to it. The author ran the benchmark and picked which grep strategy to compare against, so a reader has to run the suite themselves to see the numbers in their own environment.

How an agent reaches it

Leviathan ships two interfaces. The CLI returns short cards with the row, a citation back to the index, and a score, so a shell-using agent reads them straight out of its terminal. The MCP mode exposes four read-only tools (search, resolve_group, get, describe) that an MCP client calls directly. Both paths return the same cards. The agent does not need to know how the index works, and the index is one SQLite file that lives on the local machine.

The use the README leads with is agent memory over the long tail of a company's own data: ticket histories, log archives, order records, anything that is too long to read and too specific to a question to summarise upfront.

Source

  • Primary: elstongun, Leviathan on GitHub, Apache 2.0, created 5 October 2026, 380 stars at the time of writing

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX