>HESTIA Aggregated Data>General Rules

General Rules


The general process described here applies to all aggregations, regardless of the termType.

Finding Cycles

Aggregations are triggered for a particular product, region, and period. The region can be a country, or World. The period is 20 years in length and is fixed at specific dates (e.g., 1990-2009, 2010-2026).

For crop aggregations, Cycles can only be included if the relevant product is the primaryProduct of the Cycle. This restriction does not apply to processedFood aggregations.

Excluded Cycles

Cycles can be excluded from aggregations for two reasons. One is that they do not represent commercial practices, which means they have commercialPracticeTreatment set to FALSE. The other is that they have a blank value for the primaryProduct. This is to exclude sources that are focused on collecting measurement data and are unlikely to reflect real farms. Cycles with a 0 value for the primaryProduct are not excluded, as these are important in representing crop failures and immature periods of plantations.

Calculation of aggregated values

Number values

An aggregated value for a data item is generally a simple mean of the values in the underlying Cycles. For example, if three Cycles have the input Urea (kg N) with values 20, 30, and 90, the aggregated value is (20 + 30 + 90) / 3 = 46.7.

There are two key exceptions. One is for the primaryProduct in crop aggregations, values are adjusted to the same Dry matter before the mean is calculated. The other is if a weighting structure is used, a weighted mean is calculated. Examples of weighting structures include: organic/irrigation production, country production, and plantation phase. Further detail on these weighting structures can be found in the Crop page.

The simple mean is the case of the weighted mean where every weight is equal, so both are calculated the same way:

vˉ=ivi×wiWT\bar{v} = \sum_i v_i \times \frac{w_i}{W_T}

Where:

  • vˉ\bar{v} = the aggregated value, rounded to 4 significant figures
  • viv_i = the value contributed by the underlying Cycle or sub-aggregation ii, stored for a sub-aggregation but not for a Cycle
  • wiw_i = the weight of ii, or 1 where no weighting structure applies
  • WTW_T = the total of every wiw_i

Weights are not normalised when they are calculated, so each is divided by the total WTW_T here — which is what the percentages shown in the Cycle and Term descriptions represent. A total weight of 0 is treated as 1, so weights that all collapse to zero do not cause a division by zero.

When several Cycles share the same weight, that weight is divided between them, so a group represented by many Cycles does not count more than a group represented by few: a group's share of the result is its intended weight, not its number of uploads.

w^i=wink(i)×WT\hat{w}_i = \frac{w_i}{n_{k(i)} \times W_T}

Where:

  • w^i\hat{w}_i = the normalised weight of Cycle ii
  • wiw_i = the weight of Cycle ii
  • nk(i)n_{k(i)} = the number of Cycles sharing Cycle $i$'s weight and group
  • WTW_T = the total of every wi/nk(i)w_i / n_{k(i)}

The viv_i of the Cycles behind a sub-aggregation are not stored: a sub-aggregation can cover tens of thousands of Cycles, and copying their values into the aggregation would multiply the dataset for no information gain — only the weight each had is recorded. The handful of sub-aggregations a country aggregation combines are bounded, so their values are stored. The Cycles that went in are listed under aggregatedCycles, and their values can be read from the Cycles themselves. The number of Cycles behind a value is published as observations, and the spread as min, max, sd and the percentiles of the distribution.

Boolean values

For terms with boolean units, the aggregated value is the majority of the underlying values: it is TRUE when at least half of the underlying Cycles have the value TRUE, and FALSE otherwise. Weights are not applied to boolean values.

Completeness

Aggregations operate at the level of completeness areas. For example, all fertiliser data from Cycles where completeness.fertiliser = TRUE are aggregated together. If a Cycle has completeness.fertiliser = FALSE, the fertiliser data from that Cycle are excluded, but data from other completeness areas which are TRUE are still aggregated. This approach allows the pooling of data from many incomplete Cycles to create an overall complete dataset.

Blank values occur for two reasons: the user has uploaded a data item with a blank value, or a data item exists in one underlying Cycle but not another. For example, Cycle A has the input Urea (kg N) with value = 70, but Cycle B does not have the input Urea (kg N) despite having completeness.fertiliser = TRUE. Cycle B effectively has a blank value for Urea (kg N).

For blank values with an associated completeness area (e.g., fertiliser): if the area is complete, blank values within this area are treated as value = 0; if the area is incomplete, blank values are skipped.

This zero-filling is what allows a term reported by only part of the population to be averaged over the whole population rather than only where it occurs. It leaves no trace in the aggregated data — the zero-filled Cycles report nothing for the term — so it is the one step that cannot be reproduced from the published values alone:

vˉ=ivi×wiWR+WZ\bar{v} = \sum_i v_i \times \frac{w_i}{W_R + W_Z}

Where:

  • vˉ\bar{v} = the aggregated value
  • viv_i = the value reported for the term, stored for a sub-aggregation but not for a Cycle
  • wiw_i = the weight of a Cycle or sub-aggregation that reports the term
  • WRW_R = the total weight of those that report the term
  • WZW_Z = the total weight of those that are complete for the area but report no value

Only the reporting values contribute to the sum; the zero-filled ones enter through the denominator alone, which is what pulls the result down.

For blank values without an associated completeness area (e.g., measurement), the general rule is that blank values are skipped. There are exceptions to this for percentage values, that are detailed below.

Site

The Sites to aggregate are found by looking at the cycle.site.id of each Cycle in an aggregation. This means that if two Cycles occur on the same Site, the Site effectively appears twice in the aggregation calculations.

The infrastructure node is not included in aggregations.

Management

To aggregate Site management data, all the data items with dates that fall within the aggregation period, e.g., 2010-2026, are first averaged. If the sum of these average values is over 100 for items within the same sumIs100Group, e.g., land cover or tillage, they are rescaled so the sum is 100.

The average values are then given new start and end dates based on the aggregation period. For 2010-2026, these would be 2010-01-01 and 2026-12-31.

The same method is used to aggregate management data that falls into the previous aggregation period.

Measurement

Aggregated measurement values are calculated for each depth interval present in the underlying Cycles.

Cycle

The transport and transformation nodes are not included in the aggregated Cycle.

Inputs

Inputs are aggregated by first finding the Cycles which have the relevant completeness area set to TRUE, adding 0 values to any blank values, and then averaging the values.

Primary Product

Even if the product completeness area is marked as incomplete, the value of the primaryProduct is still aggregated, to reduce data loss and improve the yield estimate.

The economicValueShare is aggregated for all Cycles where the value of the primaryProduct is greater than 0 and completeness.product = TRUE. Cycles with 0 yield values are excluded as usually nothing is produced in these Cycles so they have no economic value. An average is taken of all available economicValueShare for each product. If the sum of these averages is over 100, they are rescaled so that the total economicValueShare of the aggregated Cycle is 100.

The example below shows the aggregation of the primary product value and economicValueShare.

Full example

cycle.idcycle.products.wheatGrain.valuecycle.products.wheatGrain.economicValueSharecycle.completeness.products
1100020TRUE
25000-TRUE
34000100TRUE
4-90FALSE
52000-FALSE
6600070FALSE
700TRUE
  • Aggregated value = mean of all Cycles with a value (including 0): (1000 + 5000 + 4000 + 2000 + 6000 + 0) / 6 = 3000
  • Aggregated economicValueShare = mean of the economicValueShare of Cycles with completeness.products = TRUE and a value greater than 0, i.e. Cycles 1 and 3: (20 + 100) / 2 = 60

Cycle 2 is excluded because it has no economicValueShare, Cycle 7 because its yield is 0, and Cycles 4, 5 and 6 because their product data are incomplete.

Because economicValueShare is a share of the Cycle's total revenue rather than a physical quantity, it also has to account for how prevalent the product is across the population. A product reported by only a low-weight part of the population holds a smaller share of the aggregation's revenue than its own sub-aggregation suggests, and is rescaled down accordingly:

EVS=evs×WPWTEVS = \overline{evs} \times \frac{W_P}{W_T}

Where:

  • EVSEVS = the aggregated economic value share, rounded to 2 decimal places
  • evs\overline{evs} = the mean of the reported economic value shares
  • WPW_P = the total weight of the sub-aggregations that report this product
  • WTW_T = the total weight of all sub-aggregations

This rescale applies to economicValueShare only. A product's value (and its min, max and sd) is deliberately left unscaled: it is the average where the product is produced, and multiplying it by the reporting fraction would turn it into a total-area intensity — collapsing, for example, a crop residue reported only by a small organic sub-aggregation from 4400 to 17.

Secondary products

For secondary products (i.e., products which are not the primaryProduct), the value and economicValueShare are only aggregated from Cycles where completeness.product = TRUE. This is because secondary products are more likely to be omitted from Cycles which do not have complete product data.

Practices

How practices are aggregated depends on the glossary they are in, their units, and whether they are part of a sumIs100Group.

Those in the Operation glossary have an associated completeness area and so are aggregated like inputs.

For practices without percentage units, e.g., Current plantation age, blank values in the underlying Cycles are skipped. This is because for a lot of these practices, it does not make sense to assume a value is 0 just because it is missing.

For practices that do have percentage units, if they are in a sumIs100Group, only Cycles where the sum across the sumIs100Group is 100 are included in the aggregation. This is to ensure the aggregated value will also sum to 100.

If a Cycle has value for at least one percentage term in a group, any other percentage terms in that group which are treated has having 0 values. This ensures that the sum of percentage terms within a group adds to 100% in the aggregation. For example, if Cycle A has tillageConventional = 100 and no value for tillageReduced or tillageNoTill, the blank values for tillageReduced and tillageNoTill are treated as 0.

For practices with percentage units that are not in a sumIs100Group or the System or Standards & Labels glossaries, e.g., Green manure incorporated, blank values in the underlying cycles are skipped. This is the same method used for practices without percentage units.

For practices with percentage units, that are in the System or Standards & Labels glossaries, blank values are treated as 0. This means that when the term is missing from a Cycle, e.g., Organic, we are assuming that its value is 0.

Emissions

Tier 1 and Tier 2

Tier 1 and 2 emissions are aggregated from the underlying Cycles, rather than being recalculated from the aggregated Cycle. This is because emissions models consider more information than is aggregated, e.g., specific dates or practices with Boolean values.

Aggregated emissions values are calculated as an average of underlying Cycles, with blank values being skipped. If the underlying Cycles have different method tiers, e.g., tier 1 and tier 2, these emissions are aggregated together and the lowest tier is used as the method tier for the aggregated emission.

Background

Background emissions are recalculated using the aggregate Cycle data as it reduces the processing time compared to aggregating the background emissions from the underlying Cycles.

The only exception to this is for background emissions associated with Seed inputs. The values are then aggregated from the underlying Cycles where completeness.seed=TRUE. This is because for Saplings, background emissions come from linked nursery Cycles. For seed, the model used to estimate the background emissions requires more information than is available in the aggregated Cycle, such as which emissions are missing.

Impact Assessment

The aggregated Impact Assessments are recalculated using the aggregated Cycle data, and the aggregated economicValueShare values.