CPG and Media Taxonomy
Track 1 - Foundations · Module 1.1
Fresh analyst
Course home
Learning ObjectivesModule 1.1 · ~30 min
Place any product question at the right rung of the CPG hierarchy (category down to SKU) and apply the sum-consistency principle as a first data check.
Read the media hierarchy (channel → platform → campaign) and classify a raw media variable into the program's harmonised buckets.
Use program vocabulary correctly: ATL vs BTL, sell-in vs sell-out, syndicated/agency/publisher/in-house data, master brand vs direct media.
The product hierarchy

Every sales dataset you will touch in this program is organised as a tree. Reading top-down: category → brand → sub-brand → sub-variant → pack size → SKU. The KT sessions walk it with a laundry example: fabric cleaning is the category, Surf is the brand, Surf Excel is a sub-brand within it, and Surf Excel Plus a sub-variant of that sub-brand. Below the sub-variant sit pack type and pack size, and at the very bottom the SKU (stock keeping unit) - the lowest grain a till can scan. Nielsen's raw files carry the same tree with a manufacturer layer on top and pack type between variant and size.

The SKU level is subtler than it looks: two 250g packs of the same variant, one carrying "+20% extra" and one "+1 sachet free", are two different SKUs with different price points and different price elasticities - recommendations genuinely differ at that grain.

The property that makes the tree usable is sum-consistency: every level must sum exactly to the level above it. All sub-brands of Surf sum to Surf; all SKUs of a pack size sum to that pack size. This is your first validation on any new sales feed - if variant volumes do not re-sum to the brand total, something is wrong with the mapping, not the market.

Geography and retail channel are cuts on top of this tree (national, regional, modern vs traditional trade), not levels within it. Which rung a cell models at is a market decision: India typically asks for reads at format level (powder / liquid / bar) while the UK asks at variant level - because that is the level UK campaigns can be tracked back to. The activation level drives the read level, which brings us to mappability below.

Two selling views to keep apart: sell-in is what the manufacturer ships into retailers (invoiced shipments); sell-out is what shoppers actually buy at the till - which is what Nielsen tracks and what this program models. Media moves shoppers, not warehouses.
ATL, BTL and the four data source types

Two pieces of client lingo you will hear in every design call. ATL (above the line) is media - TV, digital, social, search, anything a consumer gets exposed to. BTL (below the line) is promo and trade activity - discounts, displays, retailer deals. The "line" is the shop roof: above it you talk to consumers, below it you work the store.

Data itself arrives from four kinds of source, and knowing which is which tells you how much to trust it and who to chase when it is wrong:

Source typeWhat it isExamples in this program
Syndicated
Continuous multi-client trackers sold by a research houseNielsen sell-out sales, Kantar panel and brand-health data
Agency
Execution records from the client's media agencyMindshare media plans, spends, GRP-to-impressions conversions
Publisher
Platform-reported delivery data from the media ownerA video or social platform's own impression exports
In-house
The client's own systemsTrade spends, launch calendars, e-commerce sales extracts
The media hierarchy

Media has its own tree: channel → platform → campaign, with attributes such as prime/non-prime and ad length hanging off the levels rather than forming levels of their own. Raw media files arrive with 5-6 layers of hierarchy per channel; the program's presentation rule is to pick at most ~3 hierarchy columns (channel, platform, campaign) and keep the rest in the variable mapping for later pivots. A channel with a single platform simply skips the platform layer.

Channels group into three buckets, and the bucket names recur in every deliverable: traditional (TV, OOH, print, radio, cinema), digital (video, display, paid social, search, streaming audio, influencer), and e-com (retail-platform display and search). Two classification traps the trainers call out by name: CTV sits under digital, never under TV - the nomenclature line between CTV and digital video is thin, but neither belongs to the TV channel; and BVOD (broadcaster video on demand) behaves like TV even though it is streamed.

Why TV dominates the conversation: it takes 50-60% of the entire media budget in most markets and brands, so client questions concentrate there - prime vs non-prime, 15s vs 30s, which campaign within prime time. Module 1.7 makes TV the anchor of the whole media bound-setting workflow for exactly this reason.

Mappability, master brand, and the national-media rule

The single most consequential taxonomy decision in a cell is media mappability: the model's granularity is decided by the level to which media execution can be tracked back. If campaigns can be mapped to a variant, you can model at variant level; if they can only be mapped to the brand, that is master brand media - brand-level execution tested against each variant's sales as its own feature.

Two hard rules follow:

  • Never split master-brand spend across variants by sales share. Brand media genuinely impacts each variant, so it is tested as-is on each - and the sanity check is that brand media impact must always be lower than direct (variant-tagged) media impact.
  • A national campaign carries the same coefficient everywhere. National media means identical content across the country (language aside), so when it is tested in a regional model there is no reason to believe it earns a different coefficient per region - do not sum it per region as if the executions were independent. The default classification when nothing is specified for a product or region: national-level in-brand.
Harmonised buckets: the Channel Segregation dictionary

Between the raw ADS and the model sits a dictionary - the methodology workbook's Channel Segregation sheet (~60 rows in the real artifact). Its job: map every raw variable name to a harmonised name plus two bucketing dimensions - Bucket 1, the reporting family (Own Price, Competitor Price, Distribution, Macro, Seasonality/Intercept, Promo, Halo, Competitor Media, Digital, Traditional), and Bucket 2, the base/incremental split. The working rule: everything driven by paid media investment is Incremental; price, distribution, promo, macro, seasonality, competitor variables and halo are all Base. The table below is illustrative (generic names, same shape as the real dictionary):

Raw variableHarmonised nameBucket 1Bucket 2
tv_impressions_primeTVTraditionalIncremental
ooh_impressionsOOHTraditionalIncremental
paid_social_impressionsPaid SocialDigitalIncremental
search_branded_spendSearchDigitalIncremental
ecomm_search_impressionsEcomm SearchDigitalIncremental
price_per_sales_volumePriceOwn PriceBase
competitor_avg_priceComp Avg PriceCompetitor PriceBase
tdp_absTDPDistributionBase
promo_tdp_indexPromo TDPPromoBase
variantx_tv_impressions (tested on core sales)Variant X TV HaloHaloBase

Notice two things. Halo rows are media by nature but classified Base - they are cross-product effects, not the focal product's own investment. And most media rows are impressions while some (like search here) are spend - the config's use_spends mechanism covers channels where spend must proxy for impressions. Platform-level taxonomy below this dictionary is country-specific: India groups by platform type, Brazil by publisher type and ad format.

One metric convention keeps all of this comparable: convert everything to PMI - paid media impressions. A GRP-based effectiveness and an impression-based effectiveness differ by orders of magnitude; standardising to impressions is what makes CPM and cross-channel comparisons possible (module 1.8 covers the GRP conversion).
Check Yourself
A new ADS lands with a column ctv_impressions. Where does it classify?
Why: the trainers flag this as a named pitfall: CTV is part of digital regardless of the screen it plays on. BVOD is the one streamed format that behaves like TV.
Classify competitor_tv_spends_agg (top competitors' TV spend, aggregated).
Why: Incremental is reserved for the brand's own paid media. Competitor spend is a Base-side pressure variable that should hurt your volume, and its bound is set as a fraction of your own direct media contribution (module 1.7).
Classify promo_tdp_index.
Why: promo is BTL, not media - it lands in the Promo family on the Base side. It is indexed (promo TDP over total TDP) because baseline distribution would exist anyway; promo makes it more visible.
A brand-level campaign with identical content runs nationwide. In a regional model, how is it handled?
Why: the national-media rule. Splitting by share is the named pitfall; the campaign is treated as national in-brand media with one coefficient, and the master-brand check applies: its impact must stay below any direct, region-tagged media.
Sources
Authored from:
  • UL - KT (1).docx (4 Jun session, Amit Kumar Pal): product and media hierarchies, ATL/BTL, national-media rule, SKU elasticity nuance, CTV placement, TV budget share
  • UL - KT.docx (2 Jun session, Harishma S): media raw-file structure, three channel buckets, master-brand media and mappability, PMI standardisation, India format-level vs UK variant-level reads
  • MathCo Methodology Understanding_UL.xlsx sheets Channel Segregation (dictionary structure and bucket logic), Data Scope (hierarchy grid), Platform Segregration (India vs Brazil platform taxonomy)
  • Training Plan_MMX.xlsx session "MMX Data Sources understanding & Model Design" (Sujit & Arka): syndicated/agency/publisher/in-house, sell-in/sell-out
The mapping table above is illustrative with generic variable names - it mirrors the shape of the real dictionary, not its contents.