AI NewsDev toolsAnnouncement

Renmin University's Data Lab releases EvoOntology, a Claude Code and Codex plugin that teaches a data agent what the columns mean, and reports a 20-point lift on DDR-Bench

Renmin University's Data Lab has released EvoOntology, a Claude Code and Codex plugin that gives a data agent a semantic layer over its warehouse and updates that layer from the agent's own query history, under MIT with 617 stars in 21 days.

AI News

Editorial3 min read

LinkedInX
EvoOntology GitHub project card

Why it mattersA data agent that keeps rediscovering what a column means burns context on facts it already knew; a semantic layer the agent updates from its own runs is the first version of that pattern as a plugin.

A data agent is given a warehouse full of tables, asks what a column means on every question, gets it wrong twice an hour and burns context tokens on the same facts.

The repository ruc-datalab/EvoOntology, from the Data Lab at Renmin University of China, is a Claude Code and Codex plugin that stores a semantic layer over a warehouse and updates that layer from the agent's own query history. It is MIT-licensed, has 617 stars and 50 forks after 21 days, and shipped with an accompanying paper on arXiv at 2609.15779. The author team is listed as Meiduo Chong, Shaolei Zhang, Ju Fan and Xiaoyong Du.

What it does, in plain terms

The plugin builds three layers over a database. A content layer stores terms, column mappings, constraints and the evidence behind each entry. A schema layer defines what objects and relationships are allowed in the content layer. A tool layer exposes the whole thing to a coding agent through MCP tools called browse_semantics and resolve_semantics, so the agent fetches only the semantics it needs for the current question rather than loading everything.

The "evolving" part is the loop that writes back. When the agent runs a query, the result and the agent's own notes on what the query answered feed back into the content layer. A rule the agent inferred and later used successfully is marked as stronger evidence; a mapping that led to a wrong answer is downgraded.

The numbers the authors report

The project page lists three benchmarks, each against a baseline the authors describe as "ReAct without an Ontology Layer". On DDR-Bench, trajectory-wise accuracy moves from 69.5 to 89.5, a 20-point lift. On BIRD, which measures text-to-SQL execution accuracy, the score moves from 63.6 to 72.4, an 8.8-point lift. On InsightBench, which rewards an agent for the quality of the insight it extracts from a dataset, the score moves from 53.2 to 54.2, a 1-point lift. The authors describe the figures as a "four-backbone analysis subset" tested across GPT-5.5, GPT-5.6-sol, Claude Sonnet 5 and Claude Opus 4.8, and point readers to the paper for the full six-backbone results.

These are the authors' own benchmarks against their own baseline, and the paper's Tables 1, 3 and 4 are where the detail sits. The InsightBench gain is close to noise, so the headline result is DDR-Bench and the real question for a reader is whether a text-to-SQL system benefits from an ontology the agent can update, which the BIRD figure also suggests.

Who it is for

A team running a data agent on top of a warehouse, where the agent is already using Claude Code or Codex and the warehouse is big enough that the schema cannot sit in the agent's context. A one-sentence version: the agent remembers that orders.status = "C" means cancelled, not completed, because the last time it guessed completed the query gave the wrong number and the agent wrote that down.

Two limits are worth stating from the README. The 20-point lift on DDR-Bench is the headline, and 1 point on InsightBench is a baseline check that the plugin does no harm on tasks whose answer is a written insight rather than a query. The project is 21 days old, the last push is 7 days ago, and the loop that writes back to the content layer depends on the agent faithfully recording what each query answered. If the agent lies to itself, the ontology learns the lie.

Source

SourceGitHub

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX