Skip to content
Woyce Technologies
AboutTeamCareersContactStart a project →

Semantica Explained: A Knowledge Graph Layer for Explainable AI

Semantica is an open-source knowledge graph and decision-intelligence layer that sits alongside your LLM and vector store, adding provenance, deterministic reasoning, and audit trails to what your AI agents decide and why.

Semantica Explained: A Knowledge Graph Layer for Explainable AI — Woyce Technologies

Loading repository details…

——

A vector database can tell you which chunks of text were semantically closest to a query. It can't tell you why an AI system chose one vendor over another last quarter, what evidence it weighed, or whether that decision contradicts one it made a month earlier. Semantica is built specifically to close that gap — a knowledge graph and decision-intelligence layer, with W3C-standard provenance and deterministic reasoning baked in, meant to sit next to the LLM and vector store you already have rather than replace them.

That gap matters more as AI agents move from answering questions to making calls with real consequences: approving a loan, scoring a supplier, flagging a compliance issue. Regulators, auditors, and your own team will eventually ask why a particular decision was made, and a chat log plus a list of retrieved chunks is a weak answer. This explainer covers the specific gap Semantica targets, how its pipeline is built, what using it looks like in practice, where teams would realistically apply it, and the trade-offs to weigh before adding it to a production stack, including a recent security release worth knowing about.

Semantica's Knowledge Explorer, showing a live context graph, decision records, entity resolution, and an ontology hub

The Actual Gap It's Targeting

The project's own comparison table makes the positioning explicit: a vector-database-plus-RAG stack retrieves by embedding similarity and keeps no record of decisions, provenance, or reasoning. Plain LLM memory works off a token window and forgets the moment context rolls off. Neither can answer "why," only "what was retrieved." Semantica's pitch is that once an AI system is making decisions that matter — vendor selection, risk scoring, compliance calls — "what did the model retrieve" stops being the interesting question and "what did it decide, based on what evidence, and is that traceable later" becomes the one that actually matters.

Every decision in Semantica is a first-class, queryable object rather than a line in a chat log: you can trace its full causal ancestry, search for similar past decisions as precedent, and check it against declared policy rules before or after the fact. That's a meaningfully different data model than "store the conversation and hope you can grep it later."

Semantica vs. a Vector RAG Stack vs. LLM Memory

Based on the project's own positioning, the three approaches answer different questions:

CapabilityVector DB + RAGLLM context memorySemantica
How it retrievesEmbedding similarityWhatever fits in the token windowGraph queries plus vector search
Records decisionsNoNoYes, as queryable objects
ProvenanceUsually noneNoneW3C PROV-O lineage per fact
ReasoningLeft to the LLMLeft to the LLMDeterministic rules, Datalog, SPARQL
Handles contradictionsNewer chunks sit beside older onesOlder context rolls offFlags and resolves conflicts at ingestion
Answers "why?"No, only "what was retrieved"NoYes, via causal decision chains

The table isn't an argument that Semantica replaces the other two. Retrieval and language understanding still come from the stack you already run; Semantica adds the record-keeping and reasoning layer those components don't provide.

How It's Actually Built

The architecture is a real multi-stage pipeline, not a single library wearing a product name — each stage is described as an independently importable, shipping module:

Sources → Ingest → Parse → Normalize → Split → Extract → Conflict Detection → Deduplication
   → Knowledge Graph → [ Ontology · Reasoning · Provenance · Decisions ] → Enriched KG
   → Vector Store + Polyglot Graph Store (RDF & LPG) → Export / Visualize / REST · MCP · CLI

A few pieces of that pipeline are worth calling out specifically:

  • Provenance is W3C PROV-O, not a custom log format. Every fact carries source-linked lineage that exports to JSON, CSV, or RDF — a standards-based audit trail rather than something bespoke to this one tool, which matters the moment an audit trail needs to leave the system and be reviewed by someone who's never heard of Semantica.
  • Reasoning is deterministic by design. Forward chaining, a Rete network, Datalog, and SPARQL all produce fully explainable inference paths, and none of it requires an LLM call to run — the reasoning layer works the same way every time on the same inputs, which is the property you actually need if a decision has to be defensible later.
  • Conflicts get detected, not silently overwritten. When new information contradicts what's already in the graph, Semantica flags and resolves the conflict as part of ingestion, instead of the common failure mode where the newer fact just quietly clobbers the older one.
  • Storage is genuinely polyglot. Native RDF triple stores (embedded Oxigraph, Blazegraph, Apache Jena, Eclipse RDF4J) and Labeled Property Graphs (Neo4j, FalkorDB, Apache AGE, AWS Neptune) are both supported, swappable without touching application code — useful if a team already has institutional investment in one graph technology and doesn't want to migrate to adopt this.
  • Ingestion reaches into the systems data already lives in. Native connectors for Databricks (Unity Catalog and Delta Lake) and Snowflake mean tables in an existing lakehouse or warehouse become graph nodes with provenance attached directly, rather than requiring a separate export-then-reimport step.

Semantica's five-stage pipeline: ingest sources, extract entities, flag conflicts, enrich the graph with rules, provenance and decisions, then serve it.

Getting a Feel for It

Installation is a single pip install semantica, and the project ships a semantica doctor command that checks the install in a few seconds — confirming the Python version, the installed Semantica version, that a vector store backend (FAISS by default) is reachable, and that a config file exists. From there, the core workflow reads like ordinary Python rather than a DSL: you open a ContextGraph, call record_decision() with a category, scenario, reasoning, and outcome, and get back an ID you can hand to trace_decision_chain() for the full causal ancestry, find_similar_decisions() for precedent search, or check_decision_rules() to run it against a policy gate. Decisions can be causally linked to each other — a loan application decision causing an underwriting decision, which in turn influences a rate-assignment decision — so the resulting chain is a real graph, not a flat list. Access isn't limited to the Python library either: a REST API, an MCP server, and a CLI all sit on top of the same core, so a team can query decision history or run graph analytics from whatever layer fits their stack.

Causal chain of recorded decisions in Semantica: loan application, then underwriting, then rate assignment, traceable and checkable against policy.

Benefits of Semantica

The pipeline and API describe what Semantica does. These are the practical gains for teams whose AI systems make decisions that someone will later question.

Decisions you can explain later

The central benefit is that "why did the system decide this?" has an answer. Each decision is stored with its scenario, reasoning, outcome, and causal links, so a reviewer can trace it back through the evidence and earlier decisions that led to it. That is a different position from reconstructing intent out of chat logs and retrieval traces after something has gone wrong, and it's the position regulators and internal risk teams increasingly expect.

An audit trail others can read

Because provenance follows W3C PROV-O and exports to JSON, CSV, or RDF, the lineage of a fact or decision can be reviewed by people and tools outside Semantica. An auditor doesn't need to learn a proprietary log format or get access to the running system. For organisations that already produce compliance evidence, a standards-based trail slots into existing review processes more easily than something bespoke.

Reasoning that comes out the same every time

The rule engines run without an LLM call, so the same inputs always produce the same inference path. When a decision has to be defensible, reproducibility matters more than flexibility. Teams can rerun a check months later and get an identical result, which is exactly what a probabilistic model cannot promise on its own.

Contradictions surfaced instead of buried

When new information conflicts with what's already in the graph, Semantica flags it at ingestion. In a typical vector store, an updated policy simply sits beside the old one and retrieval may surface either. Catching conflicts early prevents agents from acting on stale or contradictory facts, and it shows data owners where their sources disagree.

Fits the stack you already run

Semantica sits beside an existing LLM and vector store, supports both RDF and property-graph databases, and connects directly to Databricks and Snowflake. Teams can add the accountability layer without migrating their retrieval, graph technology, or warehouse, and can reach it through Python, REST, MCP, or the CLI depending on where the need is.

Semantica Use Cases

These are the situations the project is built for, and where its extra architecture is most likely to pay for itself.

Regulated decision-making

Anywhere a decision needs to be explained after the fact, such as a loan approval, a clinical recommendation, or a compliance determination, a chat log and a list of retrieved chunks won't satisfy a reviewer. Semantica records each decision with its reasoning, evidence, and causal links, and checks it against declared policy rules. That lines up with how teams already think about explainable AI in high-stakes domains. The outcome is a trace an auditor can follow from the decision back to the facts that supported it.

GraphRAG over enterprise data

Plain vector retrieval treats documents as independent chunks, which loses the relationships between customers, products, contracts, and suppliers. Entity-aware chunking, extraction, and graph construction feeding a GraphRAG or RAG pipeline give retrieval more structure to reason over than embedding similarity alone. It is most useful when relationships between entities matter as much as the raw text, for example when an answer depends on which subsidiary signed which agreement.

Shared context for multiple agents

Once more than one agent is involved, each carrying its own disconnected memory, they contradict each other and repeat work. The project frames itself as a single shared intelligence layer that multiple agents can read from and write to, with conflicts flagged at ingestion. That is a concrete answer to the coordination problem that shows up once more than one AI agent needs to share what it has learned.

Audit and compliance reporting

Compliance teams need artefacts they can hand to someone outside the engineering team. Point-in-time graph snapshots and exports to PROV-O, SHACL, and OWL are aimed squarely at that need, producing a standards-based record rather than an internal debugging log. Teams can show what the system knew at a given date and how a decision was reached.

Precedent search for recurring decisions

Teams that make similar judgments repeatedly, such as supplier scoring or exception approvals, can search past decisions for precedent before making a new one. Finding similar historical cases, with their reasoning attached, helps keep decisions consistent over time and makes departures from precedent visible rather than accidental.

What to Weigh Before Adopting It

  • This adds real architectural surface area. A knowledge graph, an ontology and reasoning layer, and a provenance store are three additional systems to run, tune, and understand — worth it when decision traceability is a genuine requirement, and real overhead if the actual need was simpler retrieval.
  • It's still a young, fast-moving project. Created in mid-2025 with active, frequent releases — a sign of real engineering attention, but also a reason to expect the API surface to keep shifting for a while yet.
  • The most recent release is explicitly a security release, and that's worth taking seriously rather than glossing over. It closed several externally-reported vulnerabilities in the Explorer API and graph-store backends, including a critical missing-authentication gap that left every API route reachable without a credential, and a critical Cypher-injection path through unvalidated node labels and property keys. None of that is unusual for a young project moving fast on a web-facing surface — what matters is that the disclosures were public, specific, and fixed promptly, with authentication now failing closed by default instead of silently allowing anonymous access. Anyone running an earlier version with the Explorer API exposed should upgrade before relying on it.
  • Deterministic reasoning has real limits of its own. Rule-based inference is explainable precisely because it's rigid — it won't handle the fuzzy, context-dependent judgment calls an LLM makes easily. The right mental model is a division of labor: LLM for language and judgment, Semantica's reasoning layer for the parts that need to be provably consistent.
  • Getting real value requires actual ontology and rule work. SHACL constraints and policy rules don't write themselves — the payoff scales with how much domain modeling effort a team is willing to put in upfront, not something that shows up automatically on install.

Division of labor: the LLM stack handles language, judgment and similarity retrieval, while Semantica keeps decision records, provenance and deterministic rules.

Common Semantica Mistakes

The trade-offs above are properties of the tool. These are the adoption mistakes that turn them into problems.

Adopting it when retrieval was the real need

If the actual goal is better answers from documents, a knowledge graph, ontology layer, and provenance store are a lot of machinery to run. Teams that adopt Semantica because GraphRAG sounds more advanced, without a requirement to explain decisions, take on three extra systems for little benefit. The test is whether anyone will genuinely need to trace why a decision was made.

Exposing an older Explorer API

The recent security release fixed a missing-authentication gap and a Cypher-injection path. Running an earlier version with the Explorer API reachable on a network leaves decision records and graph data open to anyone who finds the endpoint. Treating the project like mature infrastructure, without pinning, reviewing, and upgrading versions, is a mistake for any young, fast-moving codebase.

Skipping the ontology and rule work

The payoff depends on modelling: entity types, relationships, SHACL constraints, and policy rules that reflect the domain. Installing the library and recording decisions with free-text reasoning captures some value, but the policy checks and reliable conflict detection only work once someone has done that modelling. Expecting results without it leads to a graph that is large but not useful.

Asking rules to make judgment calls

Deterministic reasoning is explainable because it is rigid. Trying to encode nuanced, context-dependent judgment as rules produces brittle rule sets that either miss cases or block reasonable decisions. Judgment belongs with the LLM or a person; the rule layer should hold the checks that must be provably consistent.

Recording decisions without their evidence

A decision record with an outcome but no linked evidence or causal relationships answers "what was decided" and little else. Teams that log decisions as isolated entries lose the main benefit of the data model. Linking each decision to the facts and earlier decisions it depended on is what makes the trace useful to an auditor.

Semantica Best Practices

Teams evaluating Semantica can reduce risk and get to value faster by following these practices.

  • Start with one decision type that already needs justifying. Pick a decision your organisation explains to auditors or regulators today, model it in a sandbox, and check whether the resulting trace answers the questions those reviewers actually ask.
  • Model the domain with the people who own it. Build the ontology, constraints, and policy rules together with domain experts and compliance staff, not engineers alone. Their vocabulary and rules should be what the graph encodes.
  • Keep a clear division of labour. Let the LLM handle language understanding and judgment, and reserve Semantica's rule engines for checks that must be consistent and reproducible. Document which decisions fall on which side.
  • Link every decision to its evidence. Record the facts, retrieved sources, and upstream decisions each decision relied on, using causal relationships, so traces are complete rather than isolated entries.
  • Pin versions and keep up with security releases. Track the project's releases, upgrade promptly when security fixes ship, and keep the Explorer API off public networks unless authentication is configured and verified.
  • Reuse the graph database you already run. If your organisation has investment in a particular RDF or property-graph store, use the supported backend rather than introducing another one.
  • Test exports with the people who will read them. Generate PROV-O or other exports early and have an auditor or compliance reviewer read them, then adjust what you record based on their feedback.
  • Measure the operational cost before scaling. Track the effort of running the graph store, reasoning layer, and provenance store during the pilot, including tuning and on-call time. Expand to more decision types only when the traceability benefit clearly outweighs that overhead, and retire decision types that nobody ends up querying.

Practical Takeaway

Semantica is a bet that as AI systems move from answering questions to making decisions that matter, "we retrieved the right context" stops being sufficient and "we can prove why this decision was made, and check it against policy" becomes the actual bar. For teams operating in regulated or high-stakes domains where that bar already exists on the human-decision side, it's worth evaluating specifically for the provenance and reasoning layer — not as a vector-store replacement, but as the accountability layer most RAG stacks don't have at all.

Teams building AI systems that need real decision traceability, knowledge-graph-backed retrieval, or audit-ready reasoning can get hands-on architecture help from Woyce Technologies.

FAQ

What is Semantica?

Semantica is an open-source, self-hostable knowledge graph and decision-intelligence platform that adds provenance tracking, deterministic reasoning, and ontology governance on top of an AI system's context — built for explainability and auditability rather than raw retrieval speed. In practice it sits beside your existing LLM and vector store and keeps a structured record of entities, facts, and decisions, each with source-linked lineage. That lets a team answer questions after the fact about what an AI system decided, which evidence it used, and whether the decision followed declared policy rules.

Does Semantica replace my vector database or LLM?

No — it's designed to complement an existing stack, not replace it. You keep your LLM, vector store, and agent framework; Semantica adds decision records, causal reasoning, provenance, and audit trails on top. Its own pipeline includes a vector store, FAISS by default, alongside the graph, so it can participate in retrieval, but the point is the accountability layer rather than better similarity search. Teams with an existing RAG pipeline can adopt it incrementally for the decisions that need traceability.

What does "decision provenance" mean in Semantica?

Every decision is recorded as a queryable object with a traceable causal chain — you can ask what evidence led to it, find similar historical decisions, analyze its downstream impact, and check it against defined policy rules. Provenance here follows the W3C PROV-O standard, so lineage can be exported as JSON, CSV, or RDF and reviewed outside the tool. That matters when an auditor or regulator needs to see how a decision was reached without learning a proprietary log format or having access to the running system.

Is Semantica's reasoning based on an LLM?

No — the reasoning engines (forward chaining, Rete network, Datalog, SPARQL), knowledge graph construction, and provenance layer are fully deterministic and don't require an LLM call to run, which is what makes their outputs reproducible and explainable. The same inputs always produce the same inference path. The trade-off is rigidity: rules won't handle nuanced judgment the way a language model does. The intended split is to let an LLM handle language and judgment while the rule engine handles the checks that must be provably consistent.

What graph databases does Semantica support?

Both RDF triple stores (embedded Oxigraph, Blazegraph, Apache Jena, Eclipse RDF4J) and Labeled Property Graphs (Neo4j, FalkorDB, Apache AGE, AWS Neptune), swappable without changing application code. That flexibility helps teams that already run a particular graph database avoid a migration just to adopt Semantica. It also ships native connectors for Databricks and Snowflake, so tables in an existing warehouse or lakehouse can become graph nodes with provenance attached rather than going through a separate export and reimport step.

Who should actually consider using Semantica?

Teams building AI systems in regulated or high-stakes domains — where a decision needs to be explainable, auditable, and checked against policy after the fact — are the clearest fit. Teams that just need faster or better retrieval without a compliance requirement likely don't need the added architectural surface. A good test is whether one decision type already has to be justified to auditors.

How do you install and verify a Semantica setup?

It installs with pip install semantica, and the bundled semantica doctor command checks the Python version, the installed Semantica version, vector store backend connectivity, and config file presence in a few seconds. FAISS is the default vector store backend that the check confirms. After that, the core workflow is ordinary Python: open a context graph, record decisions with their reasoning and outcome, and query them for causal chains, similar precedents, or policy checks. Expect the API to keep changing while the project is young.

Can I access Semantica without writing Python?

Yes — a REST API, an MCP server, and a CLI all sit on top of the same core library, so decision records, graph queries, and reasoning can be driven from whatever layer fits an existing stack, not just from Python code. The MCP server is the natural route for AI agents, while the REST API suits other services. If you expose the REST layer, run the current version, since a recent release fixed authentication gaps in the Explorer API.

Is it safe to expose Semantica's Explorer API on a network?

Only on the current version. The most recent release (v0.6.5) closed a critical missing-authentication vulnerability that left every Explorer API route reachable without a credential, along with a critical Cypher-injection path and several other issues. Anyone running an older version with the Explorer API exposed should upgrade immediately. Authentication now fails closed by default.

Does Semantica record relationships between decisions, or just individual decisions?

Both — decisions can be causally linked to each other with typed relationships (such as one decision causing or influencing another), so trace_decision_chain() returns a real multi-step causal graph rather than a flat list of unconnected records. For example, a loan application decision can cause an underwriting decision, which in turn influences a rate-assignment decision, and the whole chain stays traceable.

Conclusion

Most AI stacks are good at retrieving context and poor at remembering what they decided with it. Once agents make decisions with real consequences, that gap becomes a problem: nobody can reliably reconstruct why a choice was made, which evidence it used, or whether it contradicted an earlier decision or a policy rule.

Semantica's answer is a dedicated accountability layer. Decisions become queryable records with causal links, facts carry W3C PROV-O lineage, conflicts are flagged at ingestion rather than silently overwritten, and deterministic reasoning produces inference paths that come out the same every time. It works alongside an existing LLM and vector store, with REST, MCP, and CLI access for teams that don't want to work purely in Python.

The trade-offs are real. It adds a knowledge graph, a reasoning layer, and a provenance store to operate. The project is young and changing quickly, its recent security fixes mean only current versions should be exposed on a network, and its value depends on serious ontology and rule modeling work.

If you're considering it, start by choosing one decision type your organization already has to justify to auditors, model it in a sandbox, and test whether the trace answers the questions they would ask. For help designing that architecture, our AI and machine learning team can work through it with you.

WT

Woyce Technologies

AI & Engineering Team · Woyce

Woyce Technologies builds AI chatbots, LLM integrations, voice AI, and full-stack web applications for businesses in the US, UK, Europe & APAC. Based in Rajkot, Gujarat.

READY TO BUILD?

Let's build something
that actually works.

Tell us about your project. We'll be honest about whether we're the right fit — and if we are, we move fast.