Unlock the knowledge trapped in your documents
Fluree CAM scans your unstructured content — PDFs, contracts, audio, video, images, and web — extracts entities and relationships, maps them to your business vocabulary, and converts everything into structured knowledge graph triples.
The CAM Pipeline
From unstructured content to connected knowledge in seven stages.
Plug into the content systems you already run.
CAM ingests virtually any unstructured source — PDFs, audio, video, images, web — using the connectors content teams already trust. Adding a new source is configuration, not engineering.
Every format becomes a single normalized text layer.
Documents decompose into chapters, sections, headings, and embedded tables; audio is transcribed with speakers and timestamps; video tracks are processed; images go through OCR and visual classification — all without bespoke preprocessing for every format.
Every entity maps to a governed concept, not a string.
NER and LLM models identify people, organizations, products, amounts, dates, and domain concepts — then map each to the right IRI in your ontology. Auto-tagging against ITM vocabularies happens in the same pass; uncertain mentions route to review before they hit the graph.
Same word, right concept.
Ambiguous mentions resolve against surrounding context and ontological relationships, so each connects to exactly the right concept — not a near-miss with a similar string.
Documents become structurally connected, not just tagged.
CAM identifies typed relationships between entities — constrained by your ontology — so the output is RDF triples instead of disconnected tags. Knowledge inside content becomes traceable and queryable.
Vector and graph retrieval, side by side.
Embeddings get generated alongside extraction and link directly to the corresponding entity nodes. Keyword search, vector similarity, and graph traversal operate as one retrieval system instead of three siloed tools.
Extracted knowledge lands in Fluree Core as governed truth.
Triples and embeddings are written into Core with full provenance back to the source passage and extraction event — ready for AI, GraphRAG, analytics, and applications the moment they’re written.
CAM is the unstructured data on-ramp.
CAM brings documents, audio, video, and web content into the same governed semantic model that Sense, Core, and ITM share.
| Capability | Traditional Document RAG | Fluree CAM |
|---|---|---|
| What gets stored | Text chunks as vectors | Entities, relationships, and embeddings as structured triples |
| How retrieval works | Semantic similarity to the query | Graph traversal along typed relationships with optional vector similarity |
| Entity identity | None — "Apple" is just a string | Unique IRIs with semantic disambiguation |
| Relationships | Implicit in text — the model must guess | Explicit typed relationships in the graph |
| Cross-document connections | Each chunk is independent | Entities connect across all documents automatically |
| Provenance | Which chunk was retrieved | Document, entity, relationship, and extraction event preserved |
| Accuracy ceiling | ~80% | 95%+ with graph-grounded retrieval |
See CAM on your content.
Walk through your document corpus with a Fluree solutions architect.