Genloop at NeurIPS: How Enterprise AI Learns from Verified Experience

Genloop at NeurIPS: How Enterprise AI Learns from Verified Experience

Sujith P

Sujith P

Founder's Office at Genloop

Founder's Office at Genloop

Table of Contents (Add CSS Target and Preview)
Genloop is the best context infrastructure

Genloop's paper, Living Context Graphs: Continual Semantic Learning for Enterprise Data Reasoning, has been accepted to NeurIPS 2026. For our team, this is a meaningful milestone in bringing our research on enterprise AI to the wider scientific community.

NeurIPS is one of the leading research venues in AI and machine learning. Its rigorous peer-review process and highly selective program make acceptance a significant achievement. It gives us an opportunity to share the work with researchers studying how AI systems learn, reason, and use knowledge.

Genloop builds context infrastructure for AI agents: a shared foundation for agents across customers and internal teams. Product knowledge can be captured once, then adapted to each environment's systems, policies, people, and decisions. The aim is to make business context reusable and continuously maintained as deployments grow.

At the center of the paper is a question: how can an AI agent turn a verified answer into business understanding that helps with future questions?

A semantic store holds reusable business knowledge. A concept graph routes questions to validated context, making accumulated knowledge efficiently accessible.

The gap between a schema and a business answer

Consider the paper's opening example: “What were net sales last quarter?” A column named net_sales_amt provides a useful clue. It does not settle which exclusions apply or how the reporting period is defined. These conventions determine the answer, yet they may be absent from the database's explicit structure.

LCG maintains business context alongside an existing warehouse, leaving the underlying data in place. Its initial understanding is deliberately incomplete, with a process for adding and correcting knowledge through verified use.

What the system actually learns

LCG begins by profiling metadata and values, generating column definitions, inferring relationships, and grouping definitions into concepts. This provides a starting vocabulary for reasoning over the warehouse.

The semantic store distinguishes four object types: column definitions, metric definitions, scoped conventions, and concepts that connect them. These objects carry details that matter during query construction. A column definition can explain how to interpret nulls or an unusual stored format. A metric definition contains a SQL expression and the columns it references. A convention can specify a mandatory filter and the table or organizational scope in which it applies.

New knowledge is consolidated into existing semantic families. Revision links connect corrections to earlier objects; provenance distinguishes bootstrapped knowledge from learned rules and records the question that taught each rule. (Section 3; Table 4.)

How a verified answer becomes a reusable rule

In the evaluation, an initially incorrect answer enters a retry loop. The agent explores the database and generates another query; each attempt is checked against the expected result. The correction loop receives an answer-matching verdict without being shown the reference SQL. Questions that remain incorrect after the retry limit do not contribute learned rules.

Once an answer has been verified, a learning agent examines the question and execution history, identifies the context needed to solve it, and proposes a semantic addition or correction. An admission judge evaluates that proposal against the question and outcome before allowing it into the context used for future reasoning.

This creates two distinct checks: whether the answer was correct, and whether the lesson inferred from that answer is appropriate to reuse. A successful query can contain details specific to one question. Generalizing those details too widely could create a persistent error.

The retry loop corrected 80.6–86.3% of eligible questions, averaging approximately 2.9–4.1 attempts among corrected questions. This is stronger supervision than passive user feedback. (Table 7.)

A worked example: one convention across several query shapes

The paper's card_games example makes the intended transfer concrete. In this dataset, the phrase “powerful foils” corresponds to the predicate:

cardKingdomFoilId IS NOT NULL
AND cardKingdomId IS NOT NULL
cardKingdomFoilId IS NOT NULL
AND cardKingdomId IS NOT NULL
cardKingdomFoilId IS NOT NULL
AND cardKingdomId IS NOT NULL
cardKingdomFoilId IS NOT NULL
AND cardKingdomId IS NOT NULL

That convention appears in historical questions asking for counts, percentages, and the most common visual effect. A held-out question asks for a list of cards. The SQL surrounding the predicate changes substantially, while the semantic requirement remains the same.

The system needs to associate the user's wording with the stored convention, retrieve its exact expression, and combine it with the rest of the new query. The split also includes questions that require multiple conventions together: a question about borderless cards without powerful foils needs knowledge of both concepts. (Appendix C.)

Figure 2 shows how graph routing assembles context for the list question. The served neighborhood includes three column definitions, the learned predicate, and a scoped output convention. Other reachable objects are withheld: one has never been validated, while another carries negative feedback from a wrong answer.

An evaluation designed around transfer

We construct a split across three BIRD development databases: card_games, toxicology, and thrombosis_prediction. Of 499 questions, 353 form the learning history and 146 are held out. Each database is treated as an independent learning problem.

Questions are annotated with the semantic facts required by their reference answers. A held-out question is eligible only when every required fact appears in at least one different question in the learning history. This establishes that the necessary knowledge is available to learn, without exposing the target question itself during learning.

During held-out evaluation, the context state is frozen. Test questions cannot add semantic knowledge or update graph routes. The primary comparisons also withhold BIRD's per-question evidence hints and supplied database descriptions. The learner receives neither during knowledge acquisition.

This design measures transfer under a specific condition: the history contains the facts the new task requires. It does not measure the discovery of entirely unseen conventions. The held-out questions were deliberately selected to depend on knowledge absent from the schema, which explains why the initial accuracy is low. Their distribution is not a representative sample of everyday enterprise queries or the full BIRD benchmark. (Section 4.1.)

The accuracy gain comes from learned semantics

Across all 499 questions, bootstrapping raises execution accuracy from 34.1% to 38.7%. On the held-out transfer questions, the improvement is much smaller: from 13.0% to 13.7%. Profiling provides useful structure, but the selected transfer questions require conventions that structure alone does not supply.

Learning from the history changes the result:

Context supplied to the agent

Correct held-out answers

Accuracy

Raw schema

19 / 146

13.0%

Bootstrapped semantic catalog

20 / 146

13.7%

Raw schema plus five retrieved historical examples

35 / 146

24.0%

Bootstrapped catalog plus five retrieved historical examples

33 / 146

22.6%

Semantic store learned from the history

53 / 146

36.3%

Oracle: semantic store learned from all questions, including the test set

87 / 146

59.6%

Source: Table 1. All rows above omit BIRD's supplied external knowledge. Graph routing is disabled in these comparisons.

The learned store improves accuracy by 22.6 percentage points over the bootstrapped catalog, with a reported 95% confidence interval of +14.4 to +30.8 points. Improvements appear on all three databases: +28.1 points on card_games, +18.2 on toxicology, and +20.0 on thrombosis_prediction.

The retrieval baseline tests another way of using the same history. It supplies similar past questions together with their reference SQL, using a hybrid of lexical and dense retrieval. That helps substantially, reaching 24.0% in its strongest five-example configuration. The learned semantic store improves on it by a further 12.3 percentage points.

Retrieving three or ten examples does not close the pooled gap. The retrieval baseline also benefits from BM25 tuning on held-out labels. The comparison supports semantic consolidation against these tested configurations. (Tables 1 and 5.)

The most revealing failure is indiscriminate learning

When we disable the admission judge, the store grows from 410 to 488 metric definitions and from 354 to 421 conventions. Held-out accuracy falls from 53/146 to 29/146, or from 36.3% to 19.9%. The direction is consistent across all three databases. Accepting more proposed knowledge produces a less useful context layer. (Table 6.)

This ablation isolates rule admission. A separate mechanism governs graph routes after they have been used. Execution failures or negative answer feedback can remove routes from active serving, causing affected questions to fall back to semantic retrieval. The route's history is retained, and subsequent validated success can restore its influence.

Admission filters unsuitable lessons before reuse; The experiment uses an LLM admission judge. In enterprise deployment, the paper places human approval on the path that writes durable semantic updates. In Genloop's Review Center, the responsible expert reviews a proposed change and its evidence before approving it as shared context. This gives domain experts control over which lessons become reusable business knowledge. (Section 3.4; Appendix B, Table 6.)

Why the graph matters when accuracy barely changes

As the store grows, the system must find a useful subset of context for each question. Graph routes connect recurring user vocabulary to relevant concepts and their associated columns, metrics, and conventions. A full graph hit can bypass the more expensive matching pass; an incomplete hit falls back to semantic retrieval.

We test the graph's contribution by taking the final store from each incremental run and evaluating it with routing enabled and disabled:

Database

Correct answers, graph on / off

Cost per query, on / off

Median latency, on / off

card_games

22 / 23

$0.0448 / $0.1477

110 s / 373 s

toxicology

12 / 11

$0.0467 / $0.1084

181 s / 326 s

thrombosis_prediction

18 / 19

$0.1025 / $0.2755

277 s / 539 s

Source: Table 2. Each row compares the same final learned store on the same held-out questions.

Accuracy changes by at most one question per database. Disabling the graph increases cost by 2.3–3.3× and median latency by 1.8–3.4×. By the end of the streams, 63.6–78.9% of held-out questions are full graph hits.

Serving cost depends on how much context the graph can route directly. Per-query cost rises from the start to the end of all three learning streams as the semantic store expands, but individual checkpoints show declines as routing coverage improves. At the final card_games history checkpoint, cost falls 19%, from $0.0556 to $0.0448 per query, while full graph hits rise from 73.7% to 78.9%. This is one observed checkpoint in a single run. (Appendix B, Table 8.)

The paper also explores what happens when learning continues through the test questions themselves. In these continuations, per-query cost falls across all three databases: from $0.0448 to $0.0402 on card_games, $0.0467 to $0.0234 on toxicology, and $0.1025 to $0.0385 on thrombosis_prediction. These are reductions of approximately 10%, 50%, and 62%, respectively, as full graph-hit coverage reaches 91–98%. Because the graph has now learned routes for the evaluated questions themselves, these results are diagnostics using the test set, rather than evidence of the same savings on unseen questions. Whether costs eventually stabilize as coverage grows remains a hypothesis. (Appendix B, Figure 4 and Table 8.)

Learning over time, with a finite supply of new facts

Enterprise interactions arrive sequentially, so we also evaluate history in batches of approximately 20 questions. After each batch, we freeze the state and reevaluate the held-out set.

The final scores match learning from the complete history at once on card_games and toxicology; thrombosis_prediction finishes one answer lower. This supports incremental delivery of the knowledge under the tested conditions.

Learning curves plateau as later questions stop introducing facts useful to the held-out set. Repeated conventions offer limited additional benefit. The diversity of concepts in the history therefore matters when interpreting learning speed.

The study uses one run per condition, so changes of one or two answers at individual checkpoints deserve caution. Its paired-bootstrap confidence interval describes uncertainty across the evaluated questions; it does not replace repeated runs across learning orders and model randomness.

How this research fits Genloop's direction

Genloop's broader platform connects data, processes, decisions, and people in a Living Context Graph. Software companies can reuse product context while accounting for each customer's definitions and workflows. Internal teams use the same approach to give agents a shared understanding of their business. Context Intelligence Layer

The platform's learning loop puts expert review around proposed updates. Approved corrections carry into subsequent work within that customer's environment. This product direction follows the principle investigated in the paper: context needs a controlled way to learn from experience and remain useful. Self-Learning Loop

The paper tests one part of this broader direction through text-to-SQL. Its controlled transfer results support learning reusable semantics and routing them efficiently. They do not evaluate the platform's full range of workflows or customer deployments. Production introduces further questions about noisy feedback, conflicting definitions, and who should approve durable corrections. The experimental admission judge's agreement with human reviewers also remains unmeasured.

For Genloop, the connection is practical: useful business context should survive individual interactions, reach the agents that need it, and remain open to review. As that context grows, maintaining its accuracy and serving it efficiently become infrastructure problems. Living Context Graphs investigate how those responsibilities can work together.

Frequently asked questions

Does LCG fine-tune the underlying model?

No. Learning updates an external semantic store and its graph index. The experiments hold the model stack, prompts, budgets, and decoding settings fixed, allowing the comparisons to isolate changes in available context.

Does it require moving enterprise data into a graph?

No. The warehouse remains the system of record. The graph indexes context about the data, including concepts, metric definitions, and conventions. It does not become a duplicate warehouse of business records.

How does this differ from retrieving previous SQL queries?

Example retrieval supplies past questions and answers for the model to interpret again. LCG consolidates verified experience into reusable, scoped semantic rules. Its evaluation compares those rules with retrieval from the same learning history, while retaining retrieval within its own architecture.

Is a successfully executed query enough to teach the system?

Execution establishes that a query ran. The experiment additionally checks its returned answer against a reference result, then evaluates proposed lessons before admission. Production use therefore needs meaningful verification; silence from a user is a weaker signal.

Does the paper demonstrate production readiness?

It demonstrates transfer on a controlled split across three databases. Broader deployment requires further evaluation of noisy feedback, conflicting conventions, and human oversight. The study uses one run per condition and does not measure the admission judge's agreement with human reviewers.

Based on “Living Context Graphs: Continual Semantic Learning for Enterprise Data Reasoning.” Architecture: Section 3, Figure 2, and Table 4. Evaluation: Sections 4–5 and Tables 1–2. Retrieval and admission ablations: Tables 5–6. Supervision and cost analysis: Appendix B. Transfer examples: Appendix C.

Your warehouse knows more than you're getting from it.

Your warehouse knows more than you're getting from it.

Your warehouse knows more than you're getting from it.