Conference

Does SPARQL federation work in the real world? A case study over large biological SPARQL endpoints

In Proceedings of the 25th International Semantic Web Conference (2026)

Federated SPARQL querying theoretically enables data integration over distributed knowledge sources without requiring centralized data replication. In practice, federated SPARQL querying can be performed manually (i.e., the user provides SERVICE clauses that specify targets for operations) or algorithmically (i.e., an adaptive approach for automatic source assignment for operations). To date, most evaluations of SPARQL federation, both manual and algorithmic, are conducted under controlled benchmark conditions and say relatively little about how federation behaves against public endpoints in everyday use. We present a longitudinal study that executes 67 real-world federated SPARQL queries that target over 20 widely-used, large public SPARQL endpoints, using both manual and algorithmic federation approaches, at four different time points between Spring 2025 and Spring 2026. The federated queries were obtained from users of these endpoints and the majority represent complex, biologically relevant questions. We found that the current state-of-the-art algorithmic federation approaches perform substantially worse than manual approaches across all time points, and in most cases encounter errors when executing the queries tested. We also found that query execution success decreased over the time points tested for both manual and automatic methods. These findings are framed from both user and endpoint-maintainer perspectives to encourage collaborative, community-driven improvements for users and data source maintainers alike.