Technology
LA Referencia Platform
Harvesting, processing and publishing scientific metadata.
The LA Referencia platform is a modular solution for collecting records from repositories, improving their interoperability and quality, and publishing them for search and reuse. It integrates metadata harvesting, processing, management, indexing and dissemination services. It can be deployed to operate a repository network or to aggregate and publish collections.
Its architecture follows the record lifecycle: it connects to sources that implement OAI-PMH, preserves harvested metadata, runs validations and transformations, updates indexes, and provides interfaces and protocols for querying or harvesting the information again.
From repository to publication
- 01OAI-PMH sources
Source repositories and collections.
- 02Harvester
Record retrieval and tracking.
- 03Validation and transformation
Quality rules and format mappings.
- 04Indexing and publication
Search, entities and OAI-PMH provider.
The API and management interfaces allow this workflow to be configured and monitored. Actions can be scheduled and executed in a coordinated way; incremental updates reduce reprocessing when only part of the collection has changed.
Platform components
Harvester
The Harvester is the platform’s management and processing core. It connects to OAI-PMH repositories and allows sources to be organized into networks, harvesting formats and parameters to be defined, and processes to be run manually or on a schedule.
In addition to retrieving records, it coordinates subsequent stages:
- Validation: applies configurable rules to metadata and retains results and diagnostics for quality review.
- Transformation: runs mappings between formats and prepares records for publication and indexing.
- Incremental processing: detects new, modified and deleted records. When configuration and data allow, it reuses previous results and processes only changes; full runs are also supported.
- Actions and tasks: coordinates harvesting, validation, indexing and related tasks, tracking their status and results.
- Network management: groups sources and configures formats, rules, transformations and publication options for each network.
The Harvester provides a versioned management API and a web administration interface. Access is controlled through users, roles and network assignments; automated integrations can use service accounts and tokens.
Storage and traceability
The platform separates original metadata from processing results. Original records are preserved in metadata storage; catalogs and validation results retain structured information for each run. This organization allows changes to be identified, diagnostics to be consulted and results to be reused across harvests, while retaining the ability to reprocess an entire collection.
Indexing, entities and search
Transformed records can be published in indexes to make them searchable. Solr provides bibliographic indexes used for search and OAI-PMH publication. Elasticsearch or OpenSearch is used to index entities and relationships extracted from metadata. The platform also supports semantic vector indexing when an embedding generation service is configured.
VuFind can serve as a discovery interface over the bibliographic index. The availability of each interface and index depends on the deployment.
OAI-PMH provider
An independent service exposes published records through OAI-PMH 2.0. This allows other platforms and aggregators to harvest metadata from an installation. The service reads the publication index and provides an outgoing interoperability endpoint, complementing the Harvester’s harvesting from source repositories.
Web interfaces
- Administration: allows networks and processes to be configured, actions to be run, and progress, results and diagnostics to be reviewed.
- Repository dashboard: a read-only interface for monitoring available information and results, with access limited to assigned networks.
- Search: VuFind can present indexed records through a discovery interface for end users.
Persistent identifiers
Integration with dARK allows operations involving ARK identifiers to be coordinated, including reservation, preparation and reconciliation. The installation must have access to the minter service and configure the corresponding parameters.
Open source and repositories
The platform is developed across a set of Git repositories. The main platform repository brings together workspace configuration and deployment. The main components have their own repositories:
- Harvester (application) and processing library
- Administration interface and Repository dashboard
- OAI-PMH provider
- Entity model and indexing and Entity API
- Solr index configuration
- dARK/ARK integration
- OAI-PMH harvesting client
- Command-line administration tools
The main distribution is licensed under GNU AGPL v3; when reusing components, also check the license declared in each repository.
Deployment configuration: the VuFind and Dashboard interfaces, semantic indexing, OAI-PMH provider and dARK integration may require additional services, credentials or configuration in each deployment.