We use two packages from the V-Dem Institute: vdemdata,
which holds the country-year dataset, and ERT, which
identifies Episodes of Regime Transformation (get_eps(),
formerly part of vdemdata). Both are large, so everything
needed below is computed once and cached in vdem_cache.rds.
Delete that file to rebuild from a newer release.
library(dplyr)
library(tidyr)
library(ggplot2)
cache_file <- "vdem_cache.rds"
# Czechoslovakia (1918-1992) and the Czech Republic form one continuous
# V-Dem series under the code "CZE".
countries <- c(CZE = "Czechoslovakia / Czech Republic",
HUN = "Hungary",
POL = "Poland",
FRA = "France",
USA = "United States")
# Threshold grid for the episode sensitivity analysis in section 5
eps_grid <- expand_grid(start_incl = c(0.01, 0.02),
cum_incl = c(0.05, 0.10, 0.15, 0.20))
if (!file.exists(cache_file)) {
if (!requireNamespace("remotes", quietly = TRUE)) install.packages("remotes")
if (!requireNamespace("vdemdata", quietly = TRUE))
remotes::install_github("vdeminstitute/vdemdata")
if (!requireNamespace("ERT", quietly = TRUE))
remotes::install_github("vdeminstitute/ERT")
vdem_full <- vdemdata::vdem
# Five focal countries, all variables
vdem_sub <- vdem_full |>
filter(country_text_id %in% names(countries))
# All countries, a few variables, for the comparison with Polity / Freedom House
vdem_global <- vdem_full |>
select(country_name, country_text_id, year,
v2x_polyarchy, e_p_polity, e_fh_pr, e_fh_cl)
# Episodes of regime transformation: default thresholds + sensitivity grid
eps_keep <- function(e) filter(e, country_text_id %in% names(countries))
eps_default <- eps_keep(ERT::get_eps(data = vdem_full))
eps_sens <- eps_grid |>
rowwise() |>
mutate(eps = list(eps_keep(ERT::get_eps(data = vdem_full,
start_incl = start_incl,
cum_incl = cum_incl)))) |>
ungroup()
saveRDS(list(vdem_sub = vdem_sub, vdem_global = vdem_global,
eps_default = eps_default, eps_sens = eps_sens), cache_file)
rm(vdem_full)
}
cache <- readRDS(cache_file)
vdem_sub <- cache$vdem_sub |>
mutate(country = factor(countries[country_text_id], levels = countries))
vdem_global <- cache$vdem_global
eps_default <- cache$eps_default
eps_sens <- cache$eps_sens
vdem_sub |>
group_by(country) |>
summarise(first_year = min(year), last_year = max(year), n_years = n())
## # A tibble: 5 x 4
## country first_year last_year n_years
## <fct> <dbl> <dbl> <int>
## 1 Czechoslovakia / Czech Republic 1918 2025 108
## 2 Hungary 1789 2025 237
## 3 Poland 1789 2025 171
## 4 France 1789 2025 237
## 5 United States 1789 2025 237
Shared plotting objects:
indices <- c(v2x_polyarchy = "Electoral",
v2x_libdem = "Liberal",
v2x_partipdem = "Participatory",
v2x_delibdem = "Deliberative",
v2x_egaldem = "Egalitarian")
country_cols <- c("Czechoslovakia / Czech Republic" = "#1b9e77",
"Hungary" = "#d95f02",
"Poland" = "#7570b3",
"France" = "#e7298a",
"United States" = "#66a61e")
theme_set(theme_minimal(base_size = 12) +
theme(panel.grid.minor = element_blank(),
strip.text = element_text(face = "bold")))
# Long format of the five indices. complete() inserts empty years so that
# lines break across gaps in coding (e.g. Poland 1796-1808) instead of
# being drawn straight through them.
hli <- vdem_sub |>
select(country, year, all_of(names(indices))) |>
group_by(country) |>
complete(year = full_seq(year, 1)) |>
ungroup() |>
pivot_longer(all_of(names(indices)), names_to = "index", values_to = "score") |>
mutate(index = factor(indices[index], levels = indices))
V-Dem does not offer a single measure of democracy. It measures five conceptions, each scaled from 0 to 1.
Electoral democracy (v2x_polyarchy).
This is Dahl’s polyarchy and the core of all the other indices. It
combines five components: elected officials (v2x_elecoff),
clean elections (v2xel_frefair), freedom of association
(v2x_frassoc_thick), suffrage (v2x_suffr) and
freedom of expression and alternative sources of information
(v2x_freexp_altinf). The index is the average of a
multiplicative term, in which a zero on any component sinks the whole
score, and a weighted additive term, in which strengths compensate for
weaknesses.
Liberal democracy (v2x_libdem). Adds
the protection of individual and minority rights against the state and
the majority: equality before the law and individual liberties, judicial
constraints on the executive, and legislative constraints on the
executive (v2x_liberal).
Participatory democracy
(v2x_partipdem). Adds active citizen participation
beyond voting: civil society participation, direct popular votes, and
elected local and regional government (v2x_partip).
Deliberative democracy (v2xdl_delib component,
v2x_delibdem index). Adds the quality of public
reasoning: whether elites justify positions and refer to the common
good, respect counterarguments, consult widely, and whether society is
engaged in public debate.
Egalitarian democracy (v2x_egaldem).
Adds equality in the exercise of rights and freedoms: equal protection
across social groups, equal access to power, and equal distribution of
the resources needed to participate (v2x_egal).
Each of the four “thick” indices combines the electoral index with its own component using the same formula, for example for the liberal index:
\[\text{LDI} = 0.25 \cdot \text{EDI}^{1.585} + 0.25 \cdot \text{LCI} + 0.5 \cdot \text{EDI}^{1.585} \cdot \text{LCI}\]
The exponent penalises weak electoral democracy, so a country cannot score highly on any conception of democracy without holding meaningful elections.
france_events <- tibble(
year = c(1792, 1799, 1815, 1830, 1848, 1852, 1870, 1940, 1944, 1958),
label = c("First Republic", "Consulate / Empire", "Restoration", "July Monarchy",
"Second Republic", "Second Empire", "Third Republic", "Vichy",
"Liberation", "Fifth Republic"))
ggplot(filter(hli, country == "France"), aes(year, score, colour = index)) +
geom_vline(data = france_events, aes(xintercept = year),
linetype = "dotted", colour = "grey60") +
geom_text(data = france_events, aes(x = year, y = 1.02, label = label),
inherit.aes = FALSE, angle = 90, hjust = 1, vjust = -0.3,
size = 2.8, colour = "grey40") +
geom_line(linewidth = 0.7, na.rm = TRUE) +
scale_y_continuous(limits = c(0, 1.02), breaks = seq(0, 1, 0.25)) +
scale_x_continuous(breaks = seq(1800, 2025, 25)) +
scale_colour_brewer(palette = "Dark2") +
labs(title = "France: V-Dem's five democracy indices, 1789 to present",
x = NULL, y = "Index score (0-1)", colour = NULL,
caption = "Source: V-Dem Country-Year dataset.") +
theme(legend.position = "bottom")
Points for discussion:
The same five indices, now with one panel per index to compare countries. Each series starts at the country’s earliest V-Dem observation: 1789 for France, Hungary, Poland and the United States, 1918 for Czechoslovakia.
ggplot(hli, aes(year, score, colour = country)) +
geom_line(linewidth = 0.55, na.rm = TRUE) +
facet_wrap(~ index, ncol = 2) +
scale_colour_manual(values = country_cols) +
scale_y_continuous(limits = c(0, 1), breaks = seq(0, 1, 0.25)) +
scale_x_continuous(breaks = seq(1800, 2025, 50)) +
labs(title = "V-Dem democracy indices, earliest observation to present",
x = NULL, y = "Index score (0-1)", colour = NULL,
caption = "Source: V-Dem Country-Year dataset. Gaps indicate years without coding.") +
theme(legend.position = c(0.75, 0.13))
Points for discussion:
V-Dem’s expert-coded indicators are aggregated with a Bayesian item
response theory model, which estimates each coder’s reliability and
scale use. The indices therefore come with uncertainty bounds,
_codelow and _codehigh, which mark the
interval of one standard deviation around the point
estimate (roughly 68%). A 95% interval would be about twice as wide, so
overlap at 68% is a conservative test of distinguishability.
cee <- vdem_sub |>
filter(country_text_id %in% c("CZE", "HUN", "POL"), year >= 2005) |>
mutate(country = droplevels(country))
ggplot(cee, aes(year, v2x_libdem,
ymin = v2x_libdem_codelow, ymax = v2x_libdem_codehigh,
colour = country, fill = country)) +
geom_ribbon(alpha = 0.2, colour = NA) +
geom_line(linewidth = 0.8) +
geom_point(size = 1.2) +
scale_colour_manual(values = country_cols) +
scale_fill_manual(values = country_cols) +
scale_x_continuous(breaks = seq(2005, 2025, 5)) +
labs(title = "Liberal democracy index with uncertainty bounds, 2005 to present",
subtitle = "Lines are point estimates; bands span codelow to codehigh (one SD)",
x = NULL, y = "Liberal democracy index", colour = NULL, fill = NULL,
caption = "Source: V-Dem Country-Year dataset.") +
theme(legend.position = "bottom")
Which pairs are statistically distinguishable in each year? The table marks a pair as overlapping when its intervals intersect.
overlap <- function(df, a, b) {
x <- filter(df, country_text_id == a)
y <- filter(df, country_text_id == b)
inner_join(x, y, by = "year", suffix = c("_a", "_b")) |>
transmute(year,
pair = paste(a, "vs", b),
leader = ifelse(v2x_libdem_a > v2x_libdem_b, a, b),
overlap = v2x_libdem_codelow_a <= v2x_libdem_codehigh_b &
v2x_libdem_codelow_b <= v2x_libdem_codehigh_a)
}
overlaps <- bind_rows(overlap(cee, "CZE", "POL"),
overlap(cee, "CZE", "HUN"),
overlap(cee, "POL", "HUN"))
overlaps |>
mutate(cell = ifelse(overlap, paste0(leader, " (overlap)"), leader)) |>
select(year, pair, cell) |>
pivot_wider(names_from = pair, values_from = cell) |>
knitr::kable(caption = "Higher point estimate in each pair; '(overlap)' = intervals intersect")
| year | CZE vs POL | CZE vs HUN | POL vs HUN |
|---|---|---|---|
| 2005 | CZE (overlap) | CZE (overlap) | POL (overlap) |
| 2006 | CZE (overlap) | CZE (overlap) | POL (overlap) |
| 2007 | CZE (overlap) | CZE (overlap) | POL (overlap) |
| 2008 | POL (overlap) | CZE (overlap) | POL (overlap) |
| 2009 | POL (overlap) | CZE (overlap) | POL (overlap) |
| 2010 | POL (overlap) | CZE | POL |
| 2011 | POL (overlap) | CZE | POL |
| 2012 | POL (overlap) | CZE | POL |
| 2013 | POL (overlap) | CZE | POL |
| 2014 | POL (overlap) | CZE | POL |
| 2015 | POL (overlap) | CZE | POL |
| 2016 | CZE | CZE | POL |
| 2017 | CZE | CZE | POL |
| 2018 | CZE | CZE | POL |
| 2019 | CZE | CZE | POL |
| 2020 | CZE | CZE | POL |
| 2021 | CZE | CZE | POL (overlap) |
| 2022 | CZE | CZE | POL (overlap) |
| 2023 | CZE | CZE | POL |
| 2024 | CZE | CZE | POL |
| 2025 | CZE | CZE | POL |
Points for discussion:
V-Dem includes other democracy measures: Polity’s combined score
(e_p_polity, −10 to +10, available to 2018) and Freedom
House’s political rights and civil liberties ratings
(e_fh_pr, e_fh_cl, 1 = most free to 7 = least
free, from 1972). Polity’s special codes for interruption, interregnum
and transition (−66, −77, −88) are set to missing, and the two Freedom
House ratings are averaged and reversed so that higher means freer.
compare <- vdem_global |>
mutate(polity = ifelse(e_p_polity < -10, NA, e_p_polity),
fh = 8 - (e_fh_pr + e_fh_cl) / 2,
focal = country_text_id %in% names(countries))
compare_long <- compare |>
select(country_text_id, year, focal, v2x_polyarchy, polity, fh) |>
pivot_longer(c(polity, fh), names_to = "measure", values_to = "value") |>
filter(!is.na(value), !is.na(v2x_polyarchy)) |>
mutate(measure = recode(measure,
polity = "Polity (-10 to +10)",
fh = "Freedom House (1 to 7, reversed)"))
compare_long |>
group_by(measure) |>
summarise(n = n(), correlation = cor(value, v2x_polyarchy)) |>
knitr::kable(digits = 2, caption = "Correlation with V-Dem electoral democracy, all country-years")
| measure | n | correlation |
|---|---|---|
| Freedom House (1 to 7, reversed) | 8332 | 0.93 |
| Polity (-10 to +10) | 16255 | 0.86 |
ggplot(compare_long, aes(value, v2x_polyarchy)) +
geom_jitter(data = filter(compare_long, !focal), width = 0.15, height = 0,
alpha = 0.06, size = 0.6, colour = "grey30") +
geom_jitter(data = filter(compare_long, focal), width = 0.15, height = 0,
alpha = 0.6, size = 0.9, colour = "#d95f02") +
geom_smooth(method = "loess", se = FALSE, colour = "black", linewidth = 0.7) +
facet_wrap(~ measure, scales = "free_x") +
scale_y_continuous(limits = c(0, 1)) +
labs(title = "V-Dem electoral democracy against Polity and Freedom House",
subtitle = "All country-years; orange = the five focal countries",
x = "Alternative measure", y = "V-Dem electoral democracy index",
caption = "Source: V-Dem Country-Year dataset.")
How much variation does Polity’s top score hide?
compare |>
filter(!is.na(polity), !is.na(v2x_polyarchy), year >= 1946) |>
mutate(polity_band = cut(polity, c(-11, -6, 5, 9, 10),
labels = c("Autocracy (-10 to -6)", "Anocracy (-5 to 5)",
"Democracy (6 to 9)", "Full democracy (10)"))) |>
group_by(polity_band) |>
summarise(country_years = n(),
edi_min = min(v2x_polyarchy),
edi_median = median(v2x_polyarchy),
edi_max = max(v2x_polyarchy)) |>
knitr::kable(digits = 2, caption = "V-Dem electoral democracy within Polity bands, 1946-2018")
| polity_band | country_years | edi_min | edi_median | edi_max |
|---|---|---|---|---|
| Autocracy (-10 to -6) | 3394 | 0.01 | 0.14 | 0.44 |
| Anocracy (-5 to 5) | 2237 | 0.01 | 0.30 | 0.78 |
| Democracy (6 to 9) | 2108 | 0.10 | 0.62 | 0.91 |
| Full democracy (10) | 1794 | 0.29 | 0.85 | 0.92 |
And the three measures over time for the focal countries, each rescaled to 0-1:
compare |>
filter(focal, year >= 1900) |>
mutate(country = factor(countries[country_text_id], levels = countries),
`V-Dem electoral` = v2x_polyarchy,
`Polity` = (polity + 10) / 20,
`Freedom House` = (fh - 1) / 6) |>
select(country, year, `V-Dem electoral`, Polity, `Freedom House`) |>
pivot_longer(-c(country, year), names_to = "measure", values_to = "score") |>
ggplot(aes(year, score, colour = measure)) +
geom_line(linewidth = 0.6, na.rm = TRUE) +
facet_wrap(~ country, ncol = 2) +
scale_colour_manual(values = c("V-Dem electoral" = "black",
"Polity" = "#1f78b4",
"Freedom House" = "#e31a1c")) +
scale_x_continuous(breaks = seq(1900, 2025, 25)) +
labs(title = "Three democracy measures rescaled to 0-1, 1900 to present",
x = NULL, y = "Rescaled score", colour = NULL,
caption = "Polity ends in 2018; Freedom House starts in 1972.") +
theme(legend.position = c(0.75, 0.12))
Points for discussion:
The ERT algorithm (Maerz et al.) identifies episodes from annual changes in the electoral democracy index. With the default thresholds, an episode starts with an annual change of at least 0.01, becomes a manifest episode once the cumulative change reaches 0.10, and ends after five years of stagnation or a reversal (an annual change of 0.03, or a cumulative change of 0.10, in the opposite direction).
outcome_aut <- c("0" = NA, "1" = "Democratic breakdown", "2" = "Preempted breakdown",
"3" = "Diminished democracy", "4" = "Averted regression",
"5" = "Regressed autocracy", "6" = "Outcome censored")
outcome_dem <- c("0" = NA, "1" = "Democratic transition", "2" = "Preempted transition",
"3" = "Stabilised electoral autocracy", "4" = "Failed liberalisation",
"5" = "Democratic deepening", "6" = "Outcome censored")
regimes <- c("Closed autocracy", "Electoral autocracy",
"Electoral democracy", "Liberal democracy")
window <- tibble(country_text_id = c("HUN", "POL"),
from = c(2000, 2010))
ep_data <- eps_default |>
inner_join(window, by = "country_text_id") |>
filter(year >= from) |>
mutate(country = countries[country_text_id],
regime = factor(regimes[v2x_regime + 1], levels = regimes))
episode_spans <- bind_rows(
ep_data |> filter(aut_ep == 1) |>
group_by(country, id = aut_ep_id) |>
summarise(start = min(year), end = max(year), .groups = "drop") |>
mutate(type = "Autocratization"),
ep_data |> filter(dem_ep == 1) |>
group_by(country, id = dem_ep_id) |>
summarise(start = min(year), end = max(year), .groups = "drop") |>
mutate(type = "Democratization"))
ggplot(ep_data, aes(year, v2x_polyarchy)) +
geom_rect(data = episode_spans,
aes(xmin = start - 0.5, xmax = end + 0.5, ymin = -Inf, ymax = Inf, fill = type),
inherit.aes = FALSE, alpha = 0.18) +
geom_ribbon(aes(ymin = v2x_polyarchy_codelow, ymax = v2x_polyarchy_codehigh),
fill = "grey70", alpha = 0.5) +
geom_line() +
geom_point(aes(colour = regime), size = 2) +
facet_wrap(~ country, scales = "free_x") +
scale_fill_manual(values = c(Autocratization = "#e31a1c", Democratization = "#1f78b4")) +
scale_colour_manual(values = c("Closed autocracy" = "#67000d", "Electoral autocracy" = "#ef3b2c",
"Electoral democracy" = "#6baed6", "Liberal democracy" = "#08306b"),
drop = FALSE) +
labs(title = "Episodes of regime transformation: Hungary and Poland",
subtitle = "Shading = ERT episodes (default thresholds); points = Regimes of the World",
x = NULL, y = "Electoral democracy index", fill = "Episode", colour = "Regime type",
caption = "Source: V-Dem; ERT package.") +
theme(legend.position = "right")
All episodes since 1985 for the five countries:
bind_rows(
eps_default |> filter(aut_ep == 1) |>
distinct(country_text_id, id = aut_ep_id, start = aut_ep_start_year,
end = aut_ep_end_year, outcome = outcome_aut[as.character(aut_ep_outcome)],
censored = aut_ep_censored) |>
mutate(type = "Autocratization"),
eps_default |> filter(dem_ep == 1) |>
distinct(country_text_id, id = dem_ep_id, start = dem_ep_start_year,
end = dem_ep_end_year, outcome = outcome_dem[as.character(dem_ep_outcome)],
censored = dem_ep_censored) |>
mutate(type = "Democratization")) |>
filter(start >= 1985) |>
mutate(country = countries[country_text_id],
censored = ifelse(censored == 1, "ongoing", "")) |>
arrange(country, start) |>
select(country, type, start, end, outcome, censored) |>
knitr::kable(caption = "ERT episodes with default thresholds")
| country | type | start | end | outcome | censored |
|---|---|---|---|---|---|
| Czechoslovakia / Czech Republic | Democratization | 1990 | 1993 | Democratic transition | |
| Czechoslovakia / Czech Republic | Autocratization | 2013 | 2021 | Averted regression | |
| Hungary | Democratization | 1988 | 1991 | Democratic transition | |
| Hungary | Autocratization | 2006 | 2025 | Democratic breakdown | ongoing |
| Poland | Democratization | 1989 | 1992 | Democratic transition | |
| Poland | Autocratization | 2016 | 2021 | Averted regression | |
| Poland | Democratization | 2023 | 2025 | Democratic deepening | ongoing |
| United States | Autocratization | 2024 | 2025 | Outcome censored | ongoing |
The episodes above depend on researcher-chosen thresholds. Below, the start threshold (annual change needed to open an episode) and the cumulative threshold (total change needed for the episode to count) are varied, holding the termination rules at their defaults.
sens_long <- eps_sens |>
mutate(eps = lapply(eps, \(e) select(e, country_text_id, year, aut_ep, dem_ep))) |>
unnest(eps) |>
filter(year >= 1985) |>
mutate(status = case_when(aut_ep == 1 ~ "Autocratization",
dem_ep == 1 ~ "Democratization"),
country = factor(countries[country_text_id], levels = rev(countries)),
start_lab = paste("start_incl =", start_incl),
cum_lab = paste("cum_incl =", format(cum_incl, nsmall = 2)))
ggplot(filter(sens_long, !is.na(status)), aes(year, country, fill = status)) +
geom_tile(height = 0.8) +
facet_grid(start_lab ~ cum_lab) +
scale_fill_manual(values = c(Autocratization = "#e31a1c", Democratization = "#1f78b4")) +
scale_x_continuous(limits = c(1984.5, 2025.5), breaks = seq(1990, 2025, 10)) +
scale_y_discrete(drop = FALSE) +
labs(title = "Episode detection under alternative thresholds",
subtitle = "Default ERT settings: start_incl = 0.01, cum_incl = 0.10",
x = NULL, y = NULL, fill = NULL) +
theme(legend.position = "bottom", panel.grid.major.y = element_blank())
Points for discussion: