Data Processing Utilities¶
Functions for processing and transforming input datasets: convergence pathway data preparation, NGHGI-consistent RCB corrections, and RCB scenario processing.
Convergence Data Processing¶
process_emissions_data¶
fair_shares.library.utils.data.convergence.process_emissions_data ¶
process_emissions_data(
country_actual_emissions_ts: TimeseriesDataFrame,
first_allocation_year: int,
emission_category: str,
group_level: str,
unit_level: str,
ur: PlainRegistry,
) -> tuple[
DataFrame, DataFrame, dict[int, str | int | float], str
]
Process country emissions data and extract initial shares.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
country_actual_emissions_ts
|
TimeseriesDataFrame
|
Raw country emissions data. |
required |
first_allocation_year
|
int
|
Year to start allocation. |
required |
emission_category
|
str
|
Emission category to analyze. |
required |
group_level
|
str
|
Index level for grouping (e.g., 'iso3c'). |
required |
unit_level
|
str
|
Index level for units. |
required |
ur
|
PlainRegistry
|
Unit registry. |
required |
Returns:
| Type | Description |
|---|---|
tuple
|
(emissions_full_numeric, emissions_countries_full, year_to_label, start_column) |
calculate_initial_shares¶
fair_shares.library.utils.data.convergence.calculate_initial_shares ¶
calculate_initial_shares(
emissions_countries_full: DataFrame,
start_column: str,
group_level: str,
) -> tuple[Series, float]
Calculate initial emission shares from actual emissions at start year.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
emissions_countries_full
|
DataFrame
|
Country emissions data (World rows already filtered out). |
required |
start_column
|
str
|
Column label for first allocation year. |
required |
group_level
|
str
|
Index level for grouping. |
required |
Returns:
| Type | Description |
|---|---|
tuple
|
(country_totals, country_sum) where country_totals is Series of emissions by country and country_sum is the total. |
process_world_scenario_data¶
fair_shares.library.utils.data.convergence.process_world_scenario_data ¶
process_world_scenario_data(
world_scenario_emissions_ts: TimeseriesDataFrame,
first_allocation_year: int,
group_level: str,
unit_level: str,
ur: PlainRegistry,
) -> tuple[
DataFrame,
Series,
list[str],
dict[int, str | int | float],
str,
float,
]
Process world scenario emissions and calculate year fractions.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
world_scenario_emissions_ts
|
TimeseriesDataFrame
|
World emissions pathway. |
required |
first_allocation_year
|
int
|
Year to start allocation. |
required |
group_level
|
str
|
Index level for grouping. |
required |
unit_level
|
str
|
Index level for units. |
required |
ur
|
PlainRegistry
|
Unit registry. |
required |
Returns:
| Type | Description |
|---|---|
tuple
|
(emissions_world, year_fraction_of_cumulative_emissions, sorted_columns, world_year_to_label, world_start_column, world_total) |
process_population_data¶
fair_shares.library.utils.data.convergence.process_population_data ¶
process_population_data(
population_ts: TimeseriesDataFrame,
first_allocation_year: int,
group_level: str,
unit_level: str,
ur: PlainRegistry,
cumulative_start_year: int | None = None,
) -> Series
Process population data and calculate cumulative population by group.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
population_ts
|
TimeseriesDataFrame
|
Population time series. |
required |
first_allocation_year
|
int
|
Year to start allocation. |
required |
group_level
|
str
|
Index level for grouping. |
required |
unit_level
|
str
|
Index level for units. |
required |
ur
|
PlainRegistry
|
Unit registry. |
required |
cumulative_start_year
|
int | None
|
If provided, cumulative population is computed from this year instead of first_allocation_year. Must be <= first_allocation_year. This shifts entitlements toward historically populous countries when early start years (e.g. 1850) are used. |
None
|
Returns:
| Type | Description |
|---|---|
Series
|
Cumulative population by group. |
build_result_dataframe¶
fair_shares.library.utils.data.convergence.build_result_dataframe ¶
build_result_dataframe(
shares_by_group: DataFrame,
emissions_countries_index: Index,
world_time_columns: list[str],
group_level: str,
unit_level: str,
) -> DataFrame
Build final result DataFrame with proper index structure.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
shares_by_group
|
DataFrame
|
Calculated shares indexed by group. |
required |
emissions_countries_index
|
Index
|
Original emissions index for alignment. |
required |
world_time_columns
|
list[str]
|
Year columns from world scenario. |
required |
group_level
|
str
|
Index level for grouping. |
required |
unit_level
|
str
|
Index level for units. |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
Result DataFrame with proper multi-index structure. |
NGHGI Corrections¶
Functions for converting IPCC RCBs to NGHGI-consistent values following Weber et al. (2026). See Scientific Documentation for methodology.
load_world_co2_lulucf¶
fair_shares.library.utils.data.nghgi.load_world_co2_lulucf ¶
Load world-total NGHGI LULUCF CO2 timeseries from notebook-produced CSV.
Reads the world-total NGHGI-reported LULUCF CO2 values produced by notebook 107 from the active LULUCF source (Melo et al., v3.1.1 or v4.0.0). The CSV has a single row with a "source" index and string year columns. Values are in MtCO2/yr (negative = net sink).
The splice year (last year of NGHGI data) is derived dynamically from the data rather than being hardcoded.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str or Path
|
Path to |
required |
Returns:
| Type | Description |
|---|---|
tuple[DataFrame, int]
|
(nghgi_ts, splice_year) where nghgi_ts is a single-row DataFrame indexed by ["source"] with string year columns, and splice_year is the last year of NGHGI data coverage. |
Raises:
| Type | Description |
|---|---|
DataLoadingError
|
If the file does not exist or expected structure is missing |
load_bunker_timeseries¶
fair_shares.library.utils.data.nghgi.load_bunker_timeseries ¶
Load international bunker fuel CO2 timeseries from notebook-produced CSV.
Reads the intermediate CSV produced by notebook 107 (LULUCF & bunker preprocessing). The CSV has a single row with a "source" index and string year columns. Values are already in MtCO2/yr.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str or Path
|
Path to |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
Single-row DataFrame indexed by ["source"] with string year columns and values in MtCO2/yr |
Raises:
| Type | Description |
|---|---|
DataLoadingError
|
If the file does not exist or expected structure is missing |
compute_bunker_deduction¶
fair_shares.library.utils.data.nghgi.compute_bunker_deduction ¶
compute_bunker_deduction(
bunker_ts: DataFrame,
start_year: int,
net_zero_year: int,
historical_end_year: int | None = None,
) -> float
Compute cumulative international bunker fuel CO2 deduction.
Combines historical year-by-year values from the bunker timeseries with extrapolation from the last observed annual rate for years beyond the historical record.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
bunker_ts
|
DataFrame
|
Bunker fuel CO2 timeseries (from load_bunker_timeseries) in MtCO2/yr |
required |
start_year
|
int
|
Start of integration window (inclusive) |
required |
net_zero_year
|
int
|
End of integration window (inclusive) |
required |
historical_end_year
|
int
|
Last year taken from the historical timeseries (default: the last
year with a value in |
None
|
Returns:
| Type | Description |
|---|---|
float
|
Total cumulative bunker deduction in MtCO2 (always positive) |
Raises:
| Type | Description |
|---|---|
DataProcessingError
|
If historical data is insufficient for the start_year |
build_nghgi_world_co2_timeseries¶
fair_shares.library.utils.data.nghgi.build_nghgi_world_co2_timeseries ¶
Construct NGHGI-consistent world total CO2 timeseries.
For backward extension of allocation years < 2020, Weber Eq. 3 requires per-year world CO2 = fossil - bunkers + LULUCF. The world CO2-FFI series already excludes international bunkers, so the result is fossil + LULUCF, where LULUCF uses: - 2000 onwards: NGHGI LULUCF (e.g. Melo et al.) - Pre-2000: NaN (no fallback — NGHGI coverage only)
No NGHGI/BM splicing is performed. Years outside NGHGI coverage are NaN.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
fossil_ts
|
DataFrame
|
World CO2-FFI emissions timeseries (e.g. PRIMAP) in Mt CO2/yr, excluding international bunkers. Must have string year columns and a MultiIndex with (iso3c, unit, emission-category). |
required |
nghgi_ts
|
DataFrame
|
NGHGI LULUCF historical timeseries (from load_world_co2_lulucf) in MtCO2/yr. Single-row DataFrame with string year columns. |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
Single-row DataFrame with same index structure as fossil_ts but emission-category label set to "co2", containing per-year NGHGI-consistent total CO2 = fossil + LULUCF. Years outside NGHGI LULUCF coverage will be NaN. |
compute_cumulative_emissions¶
fair_shares.library.utils.data.nghgi.compute_cumulative_emissions ¶
compute_cumulative_emissions(
timeseries: DataFrame, start_year: int, end_year: int
) -> float
Integrate a single-row timeseries DataFrame over a year range.
Sums values for all years from start_year to end_year (inclusive). Missing years are skipped (not interpolated) since gap-filling is the caller's responsibility.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
timeseries
|
DataFrame
|
Single-row DataFrame with string year columns (as produced by the load_* functions in this module) |
required |
start_year
|
int
|
First year to include (inclusive) |
required |
end_year
|
int
|
Last year to include (inclusive) |
required |
Returns:
| Type | Description |
|---|---|
float
|
Cumulative sum over the requested year range |
Raises:
| Type | Description |
|---|---|
DataProcessingError
|
If no year columns fall within the requested range |
RCB Processing¶
Functions for parsing RCB scenarios and converting to allocation-ready budgets.
parse_rcb_scenario¶
fair_shares.library.utils.data.rcb.parse_rcb_scenario ¶
Parse RCB scenario string into climate assessment and quantile.
RCB scenario strings follow the format "TEMPpPROB" where TEMP is the temperature target (e.g., "1.5" or "2") and PROB is the probability as a percentage (e.g., "50" or "66").
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
scenario_string
|
str
|
RCB scenario string (e.g., "1.5p50", "2p66") |
required |
Returns:
| Type | Description |
|---|---|
tuple[str, str]
|
A tuple of (climate_assessment, quantile) as strings - climate_assessment: Temperature target with "C" suffix (e.g., "1.5C") - quantile: Probability as decimal string (e.g., "0.5") |
select_rcb_scenario_set¶
fair_shares.library.utils.data.rcb.select_rcb_scenario_set ¶
select_rcb_scenario_set(
metadata: DataFrame,
source: str,
label: str,
selection: str,
band_half_width: float = DEFAULT_PEAK_WARMING_BAND_HALF_WIDTH,
) -> DataFrame
Select the AR6 scenarios behind the deductions of one budget.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
metadata
|
DataFrame
|
AR6 scenario metadata, one row per scenario, with the columns
"Category" and "Median peak warming (MAGICCv7.5.3)". The
"peak-warming-band" rule also needs "net_zero_year" (see
|
required |
source
|
str
|
RCB source key in rcbs.yaml (e.g., "forster_2026") |
required |
label
|
str
|
Budget label (e.g., "1.7p50") |
required |
selection
|
str
|
Scenario selection rule of the source:
|
required |
band_half_width
|
float
|
Half-width of the peak-warming band in degrees C (default: 0.05) |
DEFAULT_PEAK_WARMING_BAND_HALF_WIDTH
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
The selected rows of |
Raises:
| Type | Description |
|---|---|
ConfigurationError
|
If the rule is unknown, the label has no AR6 category under the "ar6-category" rule, the label has no numeric temperature under the "peak-warming-band" rule, or the metadata lack "net_zero_year" under it |
DataProcessingError
|
If the rule selects no scenario |
convention_gap_from_baseline¶
fair_shares.library.utils.data.rcb.convention_gap_from_baseline ¶
convention_gap_from_baseline(
nghgi_lulucf: Series,
bm_direct: Series,
indirect: Series,
baseline_year: int,
nz_year: int,
splice_year: int,
) -> float
Return the NGHGI-minus-BM LULUCF gap of one scenario from the budget baseline.
The gap converts a budget in the bookkeeping-model (BM) convention to the national-inventory (NGHGI) convention. It covers the years that the published budget covers: the baseline year to the net-zero year of the scenario (Weber et al. 2026, Eqs. 2-3). The years before the baseline are not part of it; the rebase adds observed NGHGI LULUCF for those years.
- Years up to
splice_year: observed NGHGI LULUCF minus the scenario BM LULUCF (AFOLU|Direct). - Years after
splice_year: the scenario indirect flux (AFOLU|Indirect), because the direct flux cancels in the difference.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
nghgi_lulucf
|
Series
|
Observed world NGHGI LULUCF CO2, indexed by year as a string |
required |
bm_direct
|
Series
|
AFOLU|Direct of the scenario, indexed by year as a string |
required |
indirect
|
Series
|
AFOLU|Indirect of the scenario, indexed by year as a string |
required |
baseline_year
|
int
|
First year of the published budget |
required |
nz_year
|
int
|
Net-zero year of the scenario (last year of the gap) |
required |
splice_year
|
int
|
Last observed year of |
required |
Returns:
| Type | Description |
|---|---|
float
|
The cumulative gap from |
calculate_budget_from_rcb¶
fair_shares.library.utils.data.rcb.calculate_budget_from_rcb ¶
calculate_budget_from_rcb(
rcb_value: float,
allocation_year: int,
world_scenario_emissions_ts: TimeseriesDataFrame,
verbose: bool = True,
) -> float
Calculate total budget to allocate based on RCB value and allocation year.
RCB (Remaining Carbon Budget) values represent the remaining budget FROM 2020 onwards. The total budget to allocate depends on the allocation year:
- If allocation_year < 2020: Add historical emissions (allocation_year to 2019)
- If allocation_year == 2020: Use RCB directly
- If allocation_year > 2020: Subtract emissions already used (2020 to allocation_year-1)
This ensures that the budget allocation is consistent regardless of which year is chosen as the allocation starting point.
All values are in Mt * CO2. RCB values are converted from Gt to Mt during preprocessing to match the units used in world_scenario_emissions_ts.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
rcb_value
|
float
|
Remaining Carbon Budget value in Mt CO2 (from 2020 onwards) |
required |
allocation_year
|
int
|
Year when budget allocation should start |
required |
world_scenario_emissions_ts
|
TimeseriesDataFrame
|
World scenario emissions timeseries data with year columns (in Mt CO2) |
required |
verbose
|
bool
|
Whether to print detailed calculation information (default: True) |
True
|
Returns:
| Type | Description |
|---|---|
float
|
Total budget to allocate in Mt CO2 |
Raises:
| Type | Description |
|---|---|
DataProcessingError
|
If the world emissions lack any year from allocation_year to 2019 (allocation_year before 2020) or from 2020 to allocation_year - 1 (allocation_year after 2020) |
process_rcb_to_2020_baseline¶
fair_shares.library.utils.data.rcb.process_rcb_to_2020_baseline ¶
process_rcb_to_2020_baseline(
rcb_value: float,
rcb_unit: str,
rcb_baseline_year: int,
emission_category: str,
world_co2_ffi_emissions: DataFrame,
world_nghgi_lulucf_emissions: DataFrame | None = None,
world_bunker_emissions: DataFrame | None = None,
bunkers_deduction_mt: float = 0.0,
lulucf_future_deduction_mt: float = 0.0,
lulucf_nghgi_correction_mt: float = 0.0,
target_baseline_year: int = 2020,
source_name: str = "",
scenario: str = "",
verbose: bool = True,
) -> dict[str, float | str | int]
Process RCB from its original baseline year to 2020 baseline with adjustments.
This function converts RCB values from any baseline year (>= 2020) to a standardized 2020 baseline. It also applies adjustments for international bunkers and LULUCF following Weber et al. (2026).
The published RCB covers total anthropogenic CO2 from its baseline year and includes international bunkers. The world CO2-FFI series excludes them, and the bunker deduction covers 2020 to net zero. The rebase therefore adds fossil emissions and bunker emissions for 2020 to (baseline_year - 1), so each bunker year from 2020 is deducted exactly once.
The rebase always uses actual observational data (e.g. PRIMAP), never scenario projections. What enters the rebase depends on the emission category:
- co2-ffi: Rebase uses fossil CO2 only. LULUCF is omitted because it
cancels algebraically with the LULUCF decomposition term.
lulucf_future_deduction_mtsubtracts expected future (base→NZ) BM LULUCF, converting the published total-CO2 RCB to an FFI-only RCB.lulucf_nghgi_correction_mtis not used for this category. - co2: Rebase uses fossil CO2 + observed LULUCF in the national-
inventory (NGHGI) convention.
lulucf_nghgi_correction_mtapplies the NGHGI-vs-BM convention gap from the baseline year to net zero (Weber et al. 2026), which re-expresses the published budget against national-inventory accounting.lulucf_future_deduction_mtis not used for this category; the budget retains both FFI and LULUCF.
The calculation follows these steps: 1. Convert RCB from source unit to Mt * CO2e 2. If baseline_year > 2020: Add actual emissions from 2020 to (baseline_year - 1) — fossil + bunkers for co2-ffi, fossil + bunkers + NGHGI LULUCF for co2 3. Subtract bunkers deduction from 2020 to net zero (always reduces budget) 4. Apply LULUCF deduction (sign-ready from caller)
Sign convention for deduction parameters: - bunkers_deduction_mt: always positive (cumulative emissions), subtracted - lulucf_future_deduction_mt: sign-ready from caller (added directly). For co2-ffi: negated BM LULUCF → positive if LULUCF is a net source (increases fossil budget), capped at 0 if caller applies a precautionary rule. Zero for co2. - lulucf_nghgi_correction_mt: sign-ready from caller (added directly). For co2: the NGHGI-vs-BM convention gap from the baseline year to net zero (Weber et al. 2026), typically negative. Zero for co2-ffi.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
rcb_value
|
float
|
Original RCB value from the source |
required |
rcb_unit
|
str
|
Unit of the RCB value (e.g., "Gt * CO2", "Mt * CO2") |
required |
rcb_baseline_year
|
int
|
The year from which the RCB is calculated (must be >= 2020) |
required |
emission_category
|
str
|
Emission category: "co2-ffi" or "co2". Controls whether NGHGI LULUCF is included in the rebase. |
required |
world_co2_ffi_emissions
|
DataFrame
|
World-level CO2-FFI emissions timeseries with year columns (in Mt * CO2e). Excludes international bunkers. |
required |
world_nghgi_lulucf_emissions
|
DataFrame or None
|
Observed world LULUCF CO2 emissions in the national-inventory (NGHGI) convention (e.g. the Melo et al. world row), with year columns (in Mt * CO2e). Used ONLY for the co2 rebase (default: None). |
None
|
world_bunker_emissions
|
DataFrame or None
|
International bunker CO2 emissions timeseries with year columns (in Mt * CO2e). Required when baseline_year > 2020 and bunkers_deduction_mt is non-zero (default: None). |
None
|
bunkers_deduction_mt
|
float
|
Total bunker CO2 emissions from 2020 to net zero in Mt * CO2e (default: 0.0). Always positive; subtracted from budget. |
0.0
|
lulucf_future_deduction_mt
|
float
|
Projected future (2020/base → NZ) BM LULUCF adjustment in Mt * CO2e, sign-ready (default: 0.0). Non-zero for co2-ffi only, where it subtracts LULUCF to convert a total-CO2 RCB to FFI-only. |
0.0
|
lulucf_nghgi_correction_mt
|
float
|
NGHGI-vs-BM convention gap from the baseline year to net zero in Mt * CO2e, sign-ready (default: 0.0). Non-zero for co2 only, where it re-expresses the budget from bookkeeping-model to NGHGI accounting. |
0.0
|
target_baseline_year
|
int
|
Target baseline year for standardization (default: 2020) |
2020
|
source_name
|
str
|
Name of the RCB source for logging (default: "") |
''
|
scenario
|
str
|
Scenario name for logging (default: "") |
''
|
verbose
|
bool
|
Whether to print detailed calculation information (default: True) |
True
|
Returns:
| Type | Description |
|---|---|
dict
|
Dictionary containing: - 'rcb_2020_nghgi_mt': RCB adjusted to 2020 baseline in Mt * CO2e - 'rcb_original_value': Original RCB value (in source units) - 'rcb_original_unit': Original RCB unit - 'baseline_year': Original baseline year - 'rebase_total_mt': Emissions added to rebase from source year to 2020 (positive, Mt * CO2e); fossil + bunkers for co2-ffi, fossil + bunkers + observed NGHGI LULUCF for co2 - 'rebase_fossil_mt': Fossil-only component of rebase (Mt * CO2e) - 'rebase_bunkers_mt': International bunker component of rebase (Mt * CO2e) - 'rebase_lulucf_mt': Observed NGHGI LULUCF component of rebase (Mt * CO2e); only non-zero for co2 - 'deduction_bunkers_mt': Bunker fuel deduction (negative, Mt * CO2e) - 'deduction_lulucf_future_mt': projected-LULUCF deduction applied to convert total-CO2 → FFI-only. Non-zero for co2-ffi, zero for co2. - 'correction_lulucf_nghgi_mt': NGHGI-vs-BM convention correction. Non-zero for co2, zero for co2-ffi. - 'net_adjustment_mt': Total change from original to 2020 baseline (rebase + deductions + correction, Mt * CO2e) |
Raises:
| Type | Description |
|---|---|
DataProcessingError
|
If a timeseries lacks a year of the rebase, or if baseline_year > 2020
and bunkers are deducted without |
fill_rebase_years¶
fair_shares.library.utils.data.rcb.fill_rebase_years ¶
fill_rebase_years(
emissions: DataFrame,
rcb_baseline_year: int,
max_fill_years: int = DEFAULT_REBASE_FILL_MAX_YEARS,
series_name: str = "emissions",
source_name: str = "",
) -> DataFrame
Hold the last observed value of a world series over the rebase years after it.
The rebase of a budget to 2020 needs every year up to the year before the RCB baseline. When the series ends earlier, each year after its last observed year takes the last observed value. The fill is a placeholder until observed data are published, and every use raises a warning.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
emissions
|
DataFrame
|
Single-row world timeseries with string year columns |
required |
rcb_baseline_year
|
int
|
The year from which the RCB is calculated |
required |
max_fill_years
|
int
|
Largest number of years to fill (default: 1). 0 turns the fill off. |
DEFAULT_REBASE_FILL_MAX_YEARS
|
series_name
|
str
|
Name of the series for the warning (e.g., "world fossil CO2 emissions") |
'emissions'
|
source_name
|
str
|
Name of the RCB source for the warning |
''
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
The series with the filled years, or the series unchanged when no
year is missing or more than |
See Also¶
- Core Utilities: General data manipulation functions
- Math Utilities: Convergence solver and adjustments
- NGHGI Corrections (Science): Scientific methodology