Technical Methodology
How VoxPolitica collects, structures, and enriches UK parliamentary data.
This page reflects VoxPolitica's methodology document as of March 2026. All parliamentary speech displayed on VoxPolitica originates from official UK Parliament sources and is not rewritten or editorially modified.
Section 1
Overview
VoxPolitica is a UK parliamentary data platform that ingests the official record of debates in the House of Commons, Hansard, and makes that record searchable, structured, and analytically useful.
This document describes the technical architecture behind the platform: how data is sourced, how it is stored, and how three distinct automated enrichment layers, taxonomy classification, entity recognition, and statement tagging, are applied to parliamentary contributions.
This methodology is intended to provide users, researchers, and journalists with a transparent account of what VoxPolitica does and what it does not do. VoxPolitica's pipeline does not modify, editorially curate, or rewrite parliamentary speech.
Section 2
Data Sources
2.1 Hansard (Official Parliamentary Record)
VoxPolitica ingests Hansard data exclusively via the official Hansard API published by UK Parliament:
https://hansard-api.parliament.uk
The platform uses the debates search endpoint to discover debate section identifiers (GUIDs) by date, then fetches each debate section's full JSON payload. These payloads contain structured transcript entries called contributions.
VoxPolitica currently ingests House of Commons data, with architectural provision for House of Lords ingestion in future.
The completeness and accuracy of the parliamentary record shown on VoxPolitica is dependent on the completeness and accuracy of Hansard as published by UK Parliament.
2.2 Parliament Members API
MP profile and membership data is sourced from the official Parliament Members API:
Section 3
The Data Pipeline
VoxPolitica processes parliamentary data through a sequential, idempotent pipeline.
3.1 Stage 1: Debate Discovery (GUID Gathering)
The first stage queries Hansard search by calendar day from a configured start date and stores debate section GUIDs as stable identifiers for downstream processing.
3.2 Stage 2: Debate JSON Download
For each discovered GUID, VoxPolitica downloads the full debate JSON payload, including metadata and the ordered list of debate items.
3.3 Stage 3: Debate Structuring
Structured debate records are extracted from raw JSON and stored with debate-level metadata.
3.4 Stage 4: Contribution Extraction
Contributions are extracted from items where ItemType is Contribution and a valid MemberId is present.
Cleaning strips HTML tags, normalises whitespace, converts line breaks, and unescapes HTML entities. Only qualifying speech items are stored as contributions.
3.5 Stage 5: NLP Enrichment
After extraction, contributions are enriched in this order:
- Statement classification (keyword and keyphrase pattern matching)
- Taxonomy classification (embedding-based semantic classification)
- Entity recognition (neural entity linking via ReFinED)
3.6 Daily Orchestration
A daily orchestration runner executes all stages daily at around 8AM GMT.
Section 4
Taxonomy Classification
4.1 Purpose
The taxonomy classifier assigns contributions to nodes in VoxPolitica's hierarchical political topic taxonomy (for example Health, Education, Foreign Affairs, Economy and their sub-topics).
The taxonomy has three levels: Level 1 broad policy domains, Level 2 sub-domains, and Level 3 specific topics. A contribution may be classified at any level.
4.2 How the Classifier Works
The classifier is embedding-based. Contribution text and each taxonomy node representation are converted to vectors, and cosine similarity determines classification candidates.
Node representations include full hierarchical path, short description, and curated keywords embedded as semantic context rather than direct string matching rules.
Classification proceeds hierarchically. The classifier starts at Level 1 and descends only when confidence thresholds are met. Otherwise it may stop at the current level or return no classification.
| Threshold | Description |
|---|---|
| None threshold | Minimum similarity score for any classification. Below this threshold the result is unclassified ("none"). |
| Descend threshold | Minimum score required to move from one level to the next. If unmet, classification may stop at the current level. |
| Margin threshold | Minimum required gap between top candidate and runner-up. Low margins are treated as ambiguous. |
4.3 Coverage and What Classification Results Mean
The taxonomy system is conservative by design and classifies about 30% of contributions. It is a precision-oriented filter intended to identify contributions that are substantially about a topic.
Filtering by taxonomy does not return every mention of a topic. It returns contributions judged to have that topic as a dominant subject. For comprehensive mention-level coverage, use full-text search in addition to taxonomy filters.
This tradeoff is intentional: high-recall topic tagging on short parliamentary speech introduces many false positives and reduces trustworthiness.
4.4 Technical Parameters
The production classifier uses sentence-transformers/all-MiniLM-L6-v2. Thresholds are calibrated empirically against Hansard data.
Section 5
Entity Recognition
5.1 Purpose
Entity recognition identifies and links named entities mentioned in contributions (people, organisations, places, legislation, policies, and other entities) to structured knowledge base entries in Wikidata and Wikipedia.
5.2 The ReFinED Model
VoxPolitica uses a modified deployment of ReFinED (Reasoning over Fine-grained Entity Descriptions), an end-to-end entity linking model developed by Amazon Research.
Source and documentation: https://github.com/amazon-science/ReFinED
ReFinED is licensed under Apache License 2.0; VoxPolitica's modified deployment is aligned with that license.
5.3 How ReFinED Works
ReFinED performs three tasks in one model pass:
- Mention detection: identifies entity spans in text
- Entity typing: applies fine-grained Wikidata types
- Entity disambiguation: links mentions to specific knowledge-base entries
VoxPolitica runs ReFinED at full debate-section level, then maps detected spans back to contribution boundaries using character offsets captured during document assembly.
5.4 Post-Processing and the Filtered Entity Table
Raw model output is retained in full. A post-processing stage normalises spans into refined_structured_entities with one row per detected span.
A filtering stage creates the interface-facing table by requiring mention-type spans with resolved Wikidata QIDs. Excluded spans remain in underlying structured tables for auditability.
5.5 Limitations of Entity Recognition
Known limitations include:
- Indirect or shorthand references may be missed
- Ambiguous mentions can be linked to incorrect entities
- Parliamentary conventions can reduce general-domain model accuracy
- Wikidata ontology types may differ from user expectations
Entity output is best treated as an exploration index rather than a definitive or exhaustive entity ledger.
Section 6
Statement Classification
6.1 Purpose
The statement classifier identifies linguistically significant statement types within contributions, such as commitments, support/opposition, constituency references, attributed claims, and numerical or temporal claims.
6.2 How the Classifier Works
Statement classification is deterministic and rule-based. It uses compiled regular expressions against contribution text to detect predefined linguistic constructions.
Core properties:
- Deterministic output for the same input and rule version
- Transparent storage of matched span and triggered pattern
- Conservative pattern design aimed at high precision
Matching runs at sentence level; results include sentence index and sentence text for exact traceability.
6.3 Statement Types
The current rule set (safe-rules-v1.3) includes:
| Tag | What It Identifies |
|---|---|
| COMMITMENT_OR_PROMISE | Explicit government or departmental commitments, promises, or guarantees. |
| INTENTION_OR_PLAN | Government or collective intention statements such as "we intend to". |
| ACCOUNTABILITY_MECHANISM | Reporting, publication, consultation, review, or legislative follow-through commitments. |
| STATED_SUPPORT | Explicit declarations of support or endorsement. |
| STATED_OPPOSITION | Explicit declarations of opposition or rejection. |
| VOTE_INTENT | Explicit statements of voting intention. |
| ATTRIBUTED_CLAIM | Claims attributed to recognised external authorities (for example OBR, ONS, NAO, IMF, IFS, NICE). |
| CONSTITUENCY_MENTIONED | References to a speaker's constituency or constituents. |
| TIMELINE_OR_DEADLINE | Explicit temporal commitments and deadlines. |
| MONETARY_VALUE_MENTIONED | Explicit monetary amounts and currency-qualified values. |
| PERCENTAGE_CITED | Explicit percentage figures. |
6.4 Contextual Decoration
Commitment, intention, and accountability tags are decorated with additional context flags:
- Polarity (nearby negation cues)
- Conditionality (for example if/unless/subject to signals)
- Reported speech indicator (third-party attribution cues)
These flags provide context and do not alter the base tag type.
6.5 Limitations of Statement Classification
- Only programmed linguistic patterns can be detected
- Formally matching but semantically off-pattern sentences can occur
- Polarity and conditionality use proximity windows, not full syntactic parsing
- No co-reference resolution for indirect references to earlier statements
Statement tags are signposts for exploration, not verified legal or factual adjudications.
Section 7
Search Infrastructure
7.1 Full-Text Search on Contributions
VoxPolitica uses PostgreSQL native tsvector/tsquery search with English language configuration and weighted fields:
- Highest weight: speaker attribution and debate title
- Medium weight: debate location
- Standard weight: contribution text
Search is accent-insensitive via PostgreSQL unaccent. Search vectors are maintained by trigger on insert/update.
7.2 Global Search Surface
VoxPolitica maintains a global search document table that aggregates contributions, debates, MPs, and other structured entities into a unified weighted full-text surface.
This table is maintained through scheduled refresh jobs that propagate new and updated records on a regular cadence.
7.3 Entity Typeahead
Entity typeahead uses PostgreSQL pg_trgm trigram indexing on entity names to support fast prefix and approximate matching in the filtering interface.
Section 8
Data Currency and Update Frequency
VoxPolitica runs a daily pipeline. New Hansard content for the prior sitting day is usually available within hours of proceedings closing, with full enrichment typically following in later pipeline steps.
| Data Type | Update Frequency |
|---|---|
| New debate GUIDs | Daily (previous day's debates) |
| Debate JSON payloads | Daily |
| Contribution extraction | Daily |
| Statement classification | Daily (new contributions) |
| Taxonomy classification | Daily (new contributions) |
| Entity recognition | Daily (new debate sections) |
| MP member data | Daily refresh; monthly historical backfill |
| Global search surface | Daily refresh job |
A lag of up to 24-48 hours can occur between a debate occurring and full enrichment availability. VoxPolitica does not guarantee real-time or same-day coverage.
9.1 Source Data Fidelity
VoxPolitica's record of what MPs said is a direct reflection of Hansard as published by UK Parliament. VoxPolitica does not independently transcribe, interpret, or editorially modify proceedings.
9.2 Classification as Probabilistic Enrichment
Taxonomy, entity, and statement layers are automated natural-language enrichments and can produce errors. They are designed to support analysis, not replace source reading.
- Taxonomy results express modelled topic dominance, not authoritative human labeling
- Entity links may include misses, false links, or ambiguous disambiguation outcomes
- Statement tags identify linguistic constructions, not verified commitments or factual truth
9.3 Classification Results Are Not Exhaustive
Classification absence is not evidence of topic absence, entity absence, or statement absence. It means no qualifying automated match was found by the current system and rule/model settings.
9.4 Use of VoxPolitica Data
VoxPolitica enrichments are intended for research, analysis, and discovery workflows. Conclusions about MPs should be verified against full contribution text and primary sources.
Official Hansard source: hansard.parliament.uk
Section 10
Third-Party Attributions and Licensing
VoxPolitica pipeline dependencies include these third-party components and sources:
| Component | Details |
|---|---|
| ReFinED Entity Linking Model | Developed by Amazon Research, licensed under Apache 2.0. Source |
| sentence-transformers / all-MiniLM-L6-v2 | Used for taxonomy semantic embedding. Sentence-Transformers (UKP Lab), model weights published on Hugging Face under Apache 2.0. |
| Hansard API | Official UK Parliament API source. Parliamentary material reproduced under Open Parliament Licence v3.0. |
| Parliament Members API | Official UK Parliament API source under Open Parliament Licence v3.0. |
Open Parliament Licence v3.0: https://www.parliament.uk/site-information/copyright-parliament/open-parliament-licence/
Section 11
Contact and Further Information
For technical queries about VoxPolitica methodology, data sources, or classification systems, please use the VoxPolitica contact page.
This methodology reflects the platform as of March 2026 and will be revised when material changes to the architecture or pipeline occur.