Query-driven Knowledge Integration with Comunica

Ruben Taelman

TNO, 12 October 2026

Query-driven Knowledge Integration with Comunica

Ghent University – imec – IDLab, Belgium

Data is highly decentralized

Decentralized data integration is challenging for application developers

Query engines abstract access to decentralized data

Hide the complexities of reading and writing for app developers

Application ↔ Gear ↔ Globe

Image credit

SPARQL processing over centralized data

Centralization not always possible

How to query over decentralized data?

Approaches for querying over decentralized data

Federated querying: Distribute query over query APIs

Link Traversal: Exploit interlinking of documents

Queries as abstraction layer for integration

SPARQL, GraphQL, ...

Heterogeneity and Federation

Flexible and Modular Meta-Query Engine

Comunica is Open

Comunica is production-ready

Default Comunica engines

Query in the browser

Query on the command line

Learn more: https://comunica.dev/docs/query/getting_started/query_cli/

Query in an JavaScript/TypeScript app


Learn more: https://comunica.dev/docs/query/getting_started/query_app/

More on https://comunica.dev/docs/

The future of RDF and SPARQL

https://www.w3.org/TR/sparql12-query/

Federation over heterogeneous sources

Federated querying over public endpoints in practise does not work due to restrictions

Real federation findings

Link Traversal: too slow for querying over Linked Open Data

Link Traversal becomes feasible with structural assumptions

Solid pods follow structural properties

If pods expose more information, complex querying can become faster

Pods exposing shape information allows query engines to skip many links

Shape index

Pods exposing cardinality information allows engines to make better query plans

Query plan times

Client-side caching across improves performance within user sessions

SolidSessionBench

Schema alignment at query time

Partial trustworthiness of data sources

Malicious friends stating different names of others

Conclusion: Query engines are a good match for integrating decentralized data

Personal Retrieval and Integration team KNoWS / IDLab / UGent