Documentation

Technical Methodology

How VoxPolitica collects, structures, and enriches UK parliamentary data.

Version 1.0March 2026

This page reflects VoxPolitica's methodology document as of March 2026. All parliamentary speech displayed on VoxPolitica originates from official UK Parliament sources and is not rewritten or editorially modified.

Section 1

Overview

VoxPolitica is a UK parliamentary data platform that ingests the official record of debates in the House of Commons, Hansard, and makes that record searchable, structured, and analytically useful.

This document describes the technical architecture behind the platform: how data is sourced, how it is stored, and how three distinct automated enrichment layers, taxonomy classification, entity recognition, and statement tagging, are applied to parliamentary contributions.

This methodology is intended to provide users, researchers, and journalists with a transparent account of what VoxPolitica does and what it does not do. VoxPolitica's pipeline does not modify, editorially curate, or rewrite parliamentary speech.

Section 2

Data Sources

2.1 Hansard (Official Parliamentary Record)

VoxPolitica ingests Hansard data exclusively via the official Hansard API published by UK Parliament:

https://hansard-api.parliament.uk

The platform uses the debates search endpoint to discover debate section identifiers (GUIDs) by date, then fetches each debate section's full JSON payload. These payloads contain structured transcript entries called contributions.

VoxPolitica currently ingests House of Commons data, with architectural provision for House of Lords ingestion in future.

The completeness and accuracy of the parliamentary record shown on VoxPolitica is dependent on the completeness and accuracy of Hansard as published by UK Parliament.

2.2 Parliament Members API

MP profile and membership data is sourced from the official Parliament Members API:

https://members-api.parliament.uk

VoxPolitica processes parliamentary data through a sequential, idempotent pipeline.

3.1 Stage 1: Debate Discovery (GUID Gathering)

The first stage queries Hansard search by calendar day from a configured start date and stores debate section GUIDs as stable identifiers for downstream processing.

3.2 Stage 2: Debate JSON Download

For each discovered GUID, VoxPolitica downloads the full debate JSON payload, including metadata and the ordered list of debate items.

3.3 Stage 3: Debate Structuring

Structured debate records are extracted from raw JSON and stored with debate-level metadata.

3.4 Stage 4: Contribution Extraction

Contributions are extracted from items where ItemType is Contribution and a valid MemberId is present.

Cleaning strips HTML tags, normalises whitespace, converts line breaks, and unescapes HTML entities. Only qualifying speech items are stored as contributions.

3.5 Stage 5: NLP Enrichment

After extraction, contributions are enriched in this order:

  1. Statement classification (keyword and keyphrase pattern matching)
  2. Taxonomy classification (embedding-based semantic classification)
  3. Entity recognition (neural entity linking via ReFinED)

3.6 Daily Orchestration

A daily orchestration runner executes all stages daily at around 8AM GMT.

4.1 Purpose

The taxonomy classifier assigns contributions to nodes in VoxPolitica's hierarchical political topic taxonomy (for example Health, Education, Foreign Affairs, Economy and their sub-topics).

The taxonomy has three levels: Level 1 broad policy domains, Level 2 sub-domains, and Level 3 specific topics. A contribution may be classified at any level.

4.2 How the Classifier Works

The classifier is embedding-based. Contribution text and each taxonomy node representation are converted to vectors, and cosine similarity determines classification candidates.

Node representations include full hierarchical path, short description, and curated keywords embedded as semantic context rather than direct string matching rules.

Classification proceeds hierarchically. The classifier starts at Level 1 and descends only when confidence thresholds are met. Otherwise it may stop at the current level or return no classification.

ThresholdDescription
None thresholdMinimum similarity score for any classification. Below this threshold the result is unclassified ("none").
Descend thresholdMinimum score required to move from one level to the next. If unmet, classification may stop at the current level.
Margin thresholdMinimum required gap between top candidate and runner-up. Low margins are treated as ambiguous.

4.3 Coverage and What Classification Results Mean

The taxonomy system is conservative by design and classifies about 30% of contributions. It is a precision-oriented filter intended to identify contributions that are substantially about a topic.

Filtering by taxonomy does not return every mention of a topic. It returns contributions judged to have that topic as a dominant subject. For comprehensive mention-level coverage, use full-text search in addition to taxonomy filters.

This tradeoff is intentional: high-recall topic tagging on short parliamentary speech introduces many false positives and reduces trustworthiness.

4.4 Technical Parameters

The production classifier uses sentence-transformers/all-MiniLM-L6-v2. Thresholds are calibrated empirically against Hansard data.

5.1 Purpose

Entity recognition identifies and links named entities mentioned in contributions (people, organisations, places, legislation, policies, and other entities) to structured knowledge base entries in Wikidata and Wikipedia.

5.2 The ReFinED Model

VoxPolitica uses a modified deployment of ReFinED (Reasoning over Fine-grained Entity Descriptions), an end-to-end entity linking model developed by Amazon Research.

Source and documentation: https://github.com/amazon-science/ReFinED

ReFinED is licensed under Apache License 2.0; VoxPolitica's modified deployment is aligned with that license.

5.3 How ReFinED Works

ReFinED performs three tasks in one model pass:

  • Mention detection: identifies entity spans in text
  • Entity typing: applies fine-grained Wikidata types
  • Entity disambiguation: links mentions to specific knowledge-base entries

VoxPolitica runs ReFinED at full debate-section level, then maps detected spans back to contribution boundaries using character offsets captured during document assembly.

5.4 Post-Processing and the Filtered Entity Table

Raw model output is retained in full. A post-processing stage normalises spans into refined_structured_entities with one row per detected span.

A filtering stage creates the interface-facing table by requiring mention-type spans with resolved Wikidata QIDs. Excluded spans remain in underlying structured tables for auditability.

5.5 Limitations of Entity Recognition

Known limitations include:

  • Indirect or shorthand references may be missed
  • Ambiguous mentions can be linked to incorrect entities
  • Parliamentary conventions can reduce general-domain model accuracy
  • Wikidata ontology types may differ from user expectations

Entity output is best treated as an exploration index rather than a definitive or exhaustive entity ledger.

6.1 Purpose

The statement classifier identifies linguistically significant statement types within contributions, such as commitments, support/opposition, constituency references, attributed claims, and numerical or temporal claims.

6.2 How the Classifier Works

Statement classification is deterministic and rule-based. It uses compiled regular expressions against contribution text to detect predefined linguistic constructions.

Core properties:

  • Deterministic output for the same input and rule version
  • Transparent storage of matched span and triggered pattern
  • Conservative pattern design aimed at high precision

Matching runs at sentence level; results include sentence index and sentence text for exact traceability.

6.3 Statement Types

The current rule set (safe-rules-v1.3) includes:

TagWhat It Identifies
COMMITMENT_OR_PROMISEExplicit government or departmental commitments, promises, or guarantees.
INTENTION_OR_PLANGovernment or collective intention statements such as "we intend to".
ACCOUNTABILITY_MECHANISMReporting, publication, consultation, review, or legislative follow-through commitments.
STATED_SUPPORTExplicit declarations of support or endorsement.
STATED_OPPOSITIONExplicit declarations of opposition or rejection.
VOTE_INTENTExplicit statements of voting intention.
ATTRIBUTED_CLAIMClaims attributed to recognised external authorities (for example OBR, ONS, NAO, IMF, IFS, NICE).
CONSTITUENCY_MENTIONEDReferences to a speaker's constituency or constituents.
TIMELINE_OR_DEADLINEExplicit temporal commitments and deadlines.
MONETARY_VALUE_MENTIONEDExplicit monetary amounts and currency-qualified values.
PERCENTAGE_CITEDExplicit percentage figures.

6.4 Contextual Decoration

Commitment, intention, and accountability tags are decorated with additional context flags:

  • Polarity (nearby negation cues)
  • Conditionality (for example if/unless/subject to signals)
  • Reported speech indicator (third-party attribution cues)

These flags provide context and do not alter the base tag type.

6.5 Limitations of Statement Classification

  • Only programmed linguistic patterns can be detected
  • Formally matching but semantically off-pattern sentences can occur
  • Polarity and conditionality use proximity windows, not full syntactic parsing
  • No co-reference resolution for indirect references to earlier statements

Statement tags are signposts for exploration, not verified legal or factual adjudications.

7.1 Full-Text Search on Contributions

VoxPolitica uses PostgreSQL native tsvector/tsquery search with English language configuration and weighted fields:

  • Highest weight: speaker attribution and debate title
  • Medium weight: debate location
  • Standard weight: contribution text

Search is accent-insensitive via PostgreSQL unaccent. Search vectors are maintained by trigger on insert/update.

7.2 Global Search Surface

VoxPolitica maintains a global search document table that aggregates contributions, debates, MPs, and other structured entities into a unified weighted full-text surface.

This table is maintained through scheduled refresh jobs that propagate new and updated records on a regular cadence.

7.3 Entity Typeahead

Entity typeahead uses PostgreSQL pg_trgm trigram indexing on entity names to support fast prefix and approximate matching in the filtering interface.

VoxPolitica runs a daily pipeline. New Hansard content for the prior sitting day is usually available within hours of proceedings closing, with full enrichment typically following in later pipeline steps.

Data TypeUpdate Frequency
New debate GUIDsDaily (previous day's debates)
Debate JSON payloadsDaily
Contribution extractionDaily
Statement classificationDaily (new contributions)
Taxonomy classificationDaily (new contributions)
Entity recognitionDaily (new debate sections)
MP member dataDaily refresh; monthly historical backfill
Global search surfaceDaily refresh job

A lag of up to 24-48 hours can occur between a debate occurring and full enrichment availability. VoxPolitica does not guarantee real-time or same-day coverage.

9.1 Source Data Fidelity

VoxPolitica's record of what MPs said is a direct reflection of Hansard as published by UK Parliament. VoxPolitica does not independently transcribe, interpret, or editorially modify proceedings.

9.2 Classification as Probabilistic Enrichment

Taxonomy, entity, and statement layers are automated natural-language enrichments and can produce errors. They are designed to support analysis, not replace source reading.

  • Taxonomy results express modelled topic dominance, not authoritative human labeling
  • Entity links may include misses, false links, or ambiguous disambiguation outcomes
  • Statement tags identify linguistic constructions, not verified commitments or factual truth

9.3 Classification Results Are Not Exhaustive

Classification absence is not evidence of topic absence, entity absence, or statement absence. It means no qualifying automated match was found by the current system and rule/model settings.

9.4 Use of VoxPolitica Data

VoxPolitica enrichments are intended for research, analysis, and discovery workflows. Conclusions about MPs should be verified against full contribution text and primary sources.

Official Hansard source: hansard.parliament.uk

VoxPolitica pipeline dependencies include these third-party components and sources:

ComponentDetails
ReFinED Entity Linking ModelDeveloped by Amazon Research, licensed under Apache 2.0. Source
sentence-transformers / all-MiniLM-L6-v2Used for taxonomy semantic embedding. Sentence-Transformers (UKP Lab), model weights published on Hugging Face under Apache 2.0.
Hansard APIOfficial UK Parliament API source. Parliamentary material reproduced under Open Parliament Licence v3.0.
Parliament Members APIOfficial UK Parliament API source under Open Parliament Licence v3.0.

Open Parliament Licence v3.0: https://www.parliament.uk/site-information/copyright-parliament/open-parliament-licence/

For technical queries about VoxPolitica methodology, data sources, or classification systems, please use the VoxPolitica contact page.

This methodology reflects the platform as of March 2026 and will be revised when material changes to the architecture or pipeline occur.