Glossary
Shared vocabulary for data governance — definitions with links into stories, tools, and learning paths.
Filter by letter
No matches for your search.
Authoritative Source
The decided source of truth for a clear purpose — with owner and evidence. Related to SSOT; stresses purpose-binding over “one system for everything”.
Control Owner
Accountable for the effectiveness of a concrete control (e.g. access review, masking, deletion gate) — often Security/Privacy/Platform, not the Data Owner of business meaning.
FRIA
Fundamental-rights impact assessment for high-risk AI: purpose, affected people, risks, mitigations, and evidence — before go-live and on material change.
Identity Resolution
Matching and separating party/customer identities across sources (match, merge, survivorship) — with purpose limits and owners.
Minimum Viable Governance
Smallest workable governance setup for SMB/mid-market: clear decision rights, one cohort product, measurable cadence — without role sprawl and without a tool landscape as a substitute for accountability.
Operating Cadence
Recurring rhythm for reviews, approvals, intake, and adoption measurement. Cadence makes governance operational — instead of one-off project slides.
Re-identification
Risk that anonymized or synthetic data can reveal individuals again. Needs gates, tests, and release approval.
Subprocessor
Further processor in the processing chain. Governance needs inventory, contract chain, change notification, and exit/evidence paths.
Survivorship
Rules for which source attribute wins conflicts in a golden record — with owner, evidence, and review. Without survivorship, identity resolution stays unresolved conflict.
TIA
Assessment of third-country transfers: legal basis, recipients, risks, supplementary measures, and review cadence — complements SCCs, does not replace them.
Data Steward
Operational ownership for definition, quality, and use of a data domain — not documentation alone.
Data Owner
Business decision-maker for purpose, access rules, and approvals of a data product or domain.
Data Architect
Owns grain, model consistency, and contracts — so domains and marts fit together without replacing owner or steward.
Data Custodian
Technical custody of systems and storage — access upkeep, backups, runtime — usually platform/IT, not business definition.
Data Consumer
Uses data products for decisions or reports — raises quality issues but does not alone decide definition and access.
Data Product Owner
Owns product lifecycle, priorities, and consumer value — distinct from the domain Owner and the Steward.
Analytics Engineer
Builds trusted transforms, tests, and docs (often with dbt] between platform and BI.
Governance Center of Excellence
Central enablement for standards, cadence, tooling, and cross-domain escalation.
Governance Lead
Accountable for the governance operating model and sponsor evidence — not every ticket.
Data Engineer
Builds and runs pipelines, storage layers, and integration patterns — delivering reliable inputs for stewardship and analytics.
BI Developer
Implements semantics and visualization in BI tools — ideally on certified datasets, not by inventing business logic inside reports.
Technical Owner
Owns technical operability of a data product or pipeline — incidents, deployments, tooling — alongside the business Data Owner.
Executive Sponsor
Provides mandate, priority, and escalation power for governance or platform initiatives — without day-to-day stewardship work.
CDO
Leadership role for data strategy, mandate, and enterprise prioritization — not day-to-day stewardship itself.
Platform Ops
Operates orchestration, environments, and platform SLAs — keeps the pipeline machinery running.
Citizen Developer
Business user who builds reports/automations themselves — needs governed semantics, or shadow IT follows.
Stakeholder
Person or group with a stake in data decisions — needs clear R/A/C/I, or role sprawl follows.
Mandate
Formal authorization for a role or CoE to drive decisions and enforce policies.
Chief Data Officer
Executive for data strategy, governance, and value — sets priorities and decision rights, does not write every KPI personally.
Analytics Translator
Translates business questions into analytical problems and results back into decisions — bridge between business and analytics.
ML Engineer
Puts models into production: serving, monitoring, retraining pipelines — not notebook experiments alone.
Platform Engineer
Builds self-service paths (CI, environments, observability) for teams — cuts toil, does not own every dashboard.
Data Scientist
Finds patterns and models from data for decisions — experiments, but does not replace platform pipeline ownership.
Prompt Engineer
Designs prompts, evaluations, and guardrails for LLM use — focuses on behavior, not infrastructure.
AI Product Manager
Prioritizes AI features by value, risk, and measurability — ties business outcomes to model/eval metrics.
Head of Data
Leads the data team operationally (priorities, delivery, people) — often under a CDO, closer to day-to-day than pure strategy.
AI Engineer
Builds production AI features (RAG, agents, evals, serving) — bridge between research notebooks and platform.
Domain Owner
Owns a business domain including data products and decisions — not ticket prioritization alone.
Data Governance Manager
Coordinates governance program, policies, and rollout — sets the frame, does not own every dataset personally.
Analytics Lead
Leads analytics team and priorities (backlog, quality, adoption) — bridge between BI/DS and business.
DataOps Engineer
Automates and runs data pipelines (CI/CD, monitoring, recovery) — ops for data, not infra alone.
Compliance Analyst
Translates regulation into auditable requirements for data/BI — not checklists without technical context.
FinOps Analyst
Analyzes and steers cloud/data costs (showback, rightsizing) — saves money, does not replace architecture.
Information Architect
Structures information and taxonomies for findability and governance — bridge between content and data catalog.
BI Analyst
Builds reports/dashboards and translates questions into models — closer to business than pure pipeline engineering.
Reporting Analyst
Specialist for recurring, governed reports — focus on consistency and distribution, not ad-hoc exploration.
Chief Analytics Officer
Executive for analytics strategy and ROI — owns insight adoption, not just tool budgets.
Data Engineering Manager
Leads data engineering teams and prioritizes platform vs. feature work — people + delivery, less hands-on SQL.
Governance Analyst
Operationalizes policies, workflows, and metadata standards — bridge between legal/compliance and engineering.
Privacy Engineer
Implements privacy-by-design in pipelines and APIs — technical execution of consent, retention, and access.
Security Engineer
Builds secure defaults into data platforms (IAM, encryption, secrets) — not just closing pen-test tickets.
Data Product
Consumable, versioned data offering with a clear audience, contracts (SLA/SLO), ownership, and documented quality.
Grain
The smallest business statement unit of a table or mart (e.g. one order, one day, one contract].
Semantic Layer
Executable meaning layer for reusable metrics, dimensions, filter logic, and aggregation rules above physical tables. It defines how a claim is calculated; a dashboard presents the claim, and a catalog makes it discoverable.
Semantic Model
Tool-bound semantic artifact (e.g. Power BI model, Qlik app model].
Data Domain
Bounded business area of ownership, products, and decision rights.
Metric / KPI Store
Technical or logical store for governed metric definitions, versions, owners, tests, and approval status. A metric store can be part of a semantic layer, but it is not the visual presentation in a dashboard.
Data Product Certification
Explicit trust status that a product meets contract, DQ, and ownership bars.
Data Product Versioning
Explicit versions and compatibility rules for product interfaces.
Breaking Change
Interface, schema, or semantics change that invalidates existing consumers.
SLA / SLO
Service promises (SLA] and measurable reliability targets (SLO], e.g. freshness.
Reverse ETL / Activation
Push curated warehouse data back into operational tools.
Data CI/CD
Automated test and deploy of models, contracts, and quality checks.
Trusted Metrics
Metrics with a clear contract, owner, grain, and versioning — so reports consume the same truth.
Medallion Architecture
Bronze/Silver/Gold as technical zones — useful labels, not a full logical warehouse model.
KPI Contract
Agreement on definition, grain, filters, owner, and breaking-change rules for a metric — before the dashboard.
Single Source of Truth
Authoritative source / SSOT: the decided, evidenced source of truth for a clear purpose (e.g. a metric or a master record). SSOT is a governance decision with an owner and evidence — not a slogan and not a mandate to store everything in one physical system.
Master Data
Shared core entities (customer, product, location) with ownership and matching — foundation for golden records.
Reference Data
Controlled value lists and codes (country, currency, status) — small, stable, often underestimated in governance.
Data Literacy
Ability to read, question, and use data meaningfully — prerequisite for self-service without chaos.
Landing / RAW Layer
Ingest as received: preserve source payload and load identity with minimal semantic change.
Deprecation
Planned retirement of a data product or measure — including deadline, replacement, and consumer communication.
Data Product Lifecycle
Phases from intake through build, operate, versioning, to deprecation of a data product.
Subject Area
Business topic block (e.g. finance, customer) — often the scope for domains, marts, and stewardship.
Producer / Consumer
Relationship between who delivers a data product and who consumes it — contracts and SLAs clarify both sides.
MVP
Smallest usable slice of a data product — vertical slice instead of big bang, with a measurable outcome.
Conform / Standardized Layer
Technical standardization and validation before business identity integration.
Data Mart
Business-scoped analytical dataset — often the business data product layer.
Dark Data
Stored but unused/undocumented data — cost, risk, and ROT candidate.
ROT
Redundant, obsolete, or trivial data — cleanup saves cost and privacy surface.
Data Strategy
Target picture and priorities for data products, capabilities, and governance — not the tool list.
Data Roadmap
Timed sequence of outcomes and data products — tied to value streams, not tools.
Domain Ownership
Ownership follows the domain, not only a central team — core of mesh and federated governance.
First-Party Data
Customer data collected directly by the company — often most valuable, but heavily regulated.
Third-Party Data
Data from external providers — scrutinize contracts, provenance, and purpose limitation.
Zero-Party Data
Preferences/intent consciously shared by the user — keep consent and purpose clear.
Integrated Core
Shared business entities, relationships, and history across sources.
Crown Jewel Data
Highest-value/highest-sensitivity assets — maximum controls, monitoring, and ownership.
Data as a Product
Datasets are run like products: clear owner, SLA, versioning, and user promise — not an ops by-product.
Mart / Business Data Product Layer
Purpose-bound facts, dimensions, and KPI bases for a defined consumer purpose.
Data SLA
Measurable promise for a data product (freshness, availability, correctness) — without an SLA, “important” stays opinion.
Dataset Versioning
Traceable versions of a dataset (schema + content) — enables reproducibility and rollback.
Trusted Data Source
Authoritative source for a business entity — downstream may copy, not contradict without a contract.
Master Data Management
Discipline to provide consistent, versioned master data (customer, product, place) — not spreadsheet matching alone.
Semantic Contract
Binding agreement on meaning and calculation of metrics/entities — not column types alone.
Operational Data Store
Near-ops, often integrated copy for operational queries — not a historical warehouse and not a lakehouse substitute.
Metric Layer
Central definition of metrics (formula, grain, owner) — prevents per-tool “revenue” variants.
Data Product Canvas
One-page sketch of users, value, interface, quality, and owner for a data product — before building pipelines.
Domain Data Product
Data product of a business domain with a clear contract — cross-domain use only via published interfaces.
Data Sharing Agreement
Governs purpose, rights, quality, and liability when sharing data between parties — not just a SharePoint link.
Platform as a Product
Internal platform is run like a product (roadmap, UX, SLAs for teams) — not as ticket helpdesk.
Consumption Contract
Tool-specific access and interpretation of a governed product (views, semantic models, extracts].
Self-Serve Platform
Teams can provision data products/pipelines without gatekeeper tickets — guardrails instead of approval queues.
Data Mesh Principle
Guardrails for decentralized data products (domain ownership, self-serve, federated governance) — not a tool product name.
Data Contract Testing
Automated tests against the published data contract — if producer breaks, build fails, not the dashboard.
Data Portability
Data subjects can receive/transfer data in machine-readable form — technical export paths must exist.
Data Product Specification
Binding description of interface, SLA, schema, and owner — where the canvas becomes operational.
Data Democratization
Broader, safe data access for non-engineers — needs guardrails, not “everyone may access everything”.
Data Entitlement
Right to specific datasets/columns for role/person — enforceable technically, not policy PDF alone.
Data Product Roadmap
Timeline for data product features and dependencies — links business outcomes to pipeline/model work.
Data Monetization
Turns data deliberately into revenue/efficiency — needs legal basis, quality, and measurable value.
Data Usage Metrics
Metrics on queries, downloads, and dashboard use — shows value and hotspots, not just storage size.
Data Consumption Pattern
Typical read pattern (batch vs. streaming, OLAP vs. API) — drives caching, partitioning, and cost.
Data Product Discovery
Users find fitting data products in catalog/marketplace — without discovery, mesh stays architecture folklore.
Data Provenance
Tracking origin and transformations of a dataset — foundation for trust and audit.
Federated Computational Governance
Global standards enforced as automated checks/policies — governance as code, not CoE slides alone.
Data Provenance Chain
Continuous chain of sources, jobs, and releases to the report — stronger than isolated lineage snapshots.
Lakehouse
Open table formats plus warehouse-style governance on shared lake storage.
Data Mesh
Domain-owned data products with federated governance, not a tooling checklist.
Simplest Viable Architecture
Least unnecessary complexity that still meets the requirement — not fewest boxes.
Modern Data Warehouse
Governed path from source to products and semantics — not just a cloud vendor rename.
Greenfield
Build-from-scratch warehouse without strangling a legacy estate.
Brownfield
Modernize an existing warehouse or estate in place.
Strangler Pattern
Incrementally replace legacy paths by routing new traffic to the modern stack.
Vertical Slice
End-to-end thin path (source → product → consumer] before horizontal platform sprawl.
Hybrid Cloud
On-prem and cloud coexistence as a deliberate architecture, not a temporary defect.
ELT
Load first, transform in the platform (typical dbt pattern].
Orchestration
Scheduling and dependency management across pipelines and quality gates.
dbt
SQL-first transform, test, and documentation framework in the warehouse/lakehouse.
Microsoft Fabric
Unified analytics SaaS (OneLake, Lakehouse/Warehouse, Power BI, Purview touchpoints].
OneLake
Fabric’s single logical lake storage plane for Lakehouse and Warehouse assets.
Apache Spark
Distributed compute for batch/streaming on lakes and warehouses — often behind notebooks and lakehouse engines.
Notebook
Interactive compute surface for exploration and jobs — production-ready only with versioning, tests, and ownership.
SQL Warehouse
SQL compute endpoint on lake/warehouse data — separates storage from elastic query compute.
Warehouse-Native Transformation
Transformations run inside the warehouse/lakehouse (ELT), not on a separate ETL server.
Parallel Run
Old and new pipeline running together — compare before cutover to catch regressions and trust gaps.
FinOps
Cost accountability for cloud/compute use — tags, budgets, and idle cleanup belong in platform governance.
dbt Seed
Versioned CSV/flat data in dbt — typical for reference data and small lookups.
dbt Snapshot
dbt mechanism for historized snapshots of source tables — related to SCD2/as-was history.
dbt Exposure
Declared downstream use (dashboard, ML, app) in dbt — makes impact and ownership visible.
Delta Lake
ACID table format commonly underlying lakehouse cores and DQ result stores.
dbt Macro
Reusable Jinja logic in dbt — e.g. for RAW generation, tests, and naming conventions.
Materialization
How dbt builds a model (view, table, incremental, ephemeral) — drives cost, latency, and downstream contracts.
Incremental Model
dbt materialization that processes only new/changed rows — needs clear unique keys and a late-arrival strategy.
Ephemeral Model
dbt model without a persisted object — inlined as a CTE into downstream models.
ETL
Transform before load — classic pattern; often replaced or hybridized by ELT in the warehouse today.
Transformation as Code
Transformations versioned in the repo with review and CI — instead of click-ETL and undocumented jobs.
Cutover
Moment consumers switch from old to new stack — needs parallel run, rollback, and clear contracts.
Parquet
Columnar file format for analytical workloads — standard in lakes and many warehouse exports.
Apache Iceberg
Open table format for lakehouses — ACID, schema evolution, and time travel on object storage.
Apache Ranger
Policy engine for fine-grained access control on Hadoop/lakehouse workloads — often paired with Atlas for audit and lineage.
QVD
Qlik’s optimized columnar extract format for reusable load layers.
Hive Metastore
Central metadata service for tables, partitions and schemas in Hadoop/lakehouse stacks — foundation for Spark, Hive and Ranger policies.
ODS
Integrated operational staging close to source systems — often a precursor to the warehouse, not a mart replacement.
Apache Atlas
Open-source governance and lineage catalog for the Hadoop ecosystem — classification, tags and audit events for Ranger-protected assets.
Data Lake
Storage for raw and semi-structured data at scale — without warehouse semantics; often a precursor or part of a lakehouse.
Apache NiFi
Visual data movement and integration framework — flows need ownership, provenance and access control like any pipeline.
Data Vault
Modeling pattern with hubs, links, and satellites for auditable historization — alternative/complement to a dimensional core.
AWS Lake Formation
AWS service for lake governance — central permissions, LF tags and data location management across S3, Glue and analytics engines.
System of Record
The authoritative operational system for an entity or attributes — not the same as an analytical single source of truth.
AWS Glue
Serverless ETL and metadata catalog in AWS — crawlers, jobs and catalog tables as the base for Lake Formation and Redshift/Spectrum.
Data Fabric
Architecture approach with metadata, virtualization, and integration across distributed sources — often a marketing umbrella beside mesh/lakehouse.
Apache Kafka
Distributed event streaming — topics, schemas and ACLs need governance analogous to batch pipelines (registry, ownership, PII tags).
Logical Architecture
Layers, contracts, and responsibilities independent of specific tooling — the foundation before physical stack choice.
Kafka ACL
Access control on Kafka resources (topics, consumer groups, cluster) — keep authority and enforcement separate from schema and contract governance.
Streaming
Continuous event/record processing — only where latency and cost justify the complexity.
Confluent
Commercial Kafka ecosystem (Platform/Cloud) with Schema Registry, Connect and governance features — same concerns as open-source Kafka, often with a stronger managed control plane.
Near Real-Time
Latency of seconds to a few minutes — costlier and more complex than batch; only where decisions truly need it.
Batch Processing
Scheduled, periodic processing of large volumes — the default for most warehouse and mart loads.
Apache Hadoop
Distributed storage and compute ecosystem (HDFS, YARN and related services) — foundation for many on-prem/hybrid lakehouse landscapes including Cloudera CDP.
Dimensional Modeling
Facts and dimensions organized for analytics grain and reuse (Kimball-style].
HDFS
Distributed file system in the Hadoop stack — blocks, replication and paths are typical Ranger/audit resources.
Jinja
Templating language behind dbt macros and models — powerful, but unreadable without standards.
Backfill
Loading historical periods after the fact — needs idempotency, watermarks, and clear cutover rules.
YARN
Cluster resource manager for Hadoop workloads — queues and compute limits belong to the operating boundary alongside data access.
Apache Hive
SQL engine and table model on the lake — policies and ownership often attach to Hive tables while the Hive Metastore holds technical truth.
Watermark
Progress marker for incremental loads (timestamp/id) — set wrong, you get gaps or duplicates.
Apache Ozone
Object storage for Hadoop/CDP landscapes — scalable alternative or complement to HDFS with its own resource boundaries for policies.
Idempotent Load
Re-running the same job does not change the result uncontrollably — mandatory for retries and backfills.
Full Load
Full reload of a table/partition — simple, expensive; often only for bootstrap or repair.
Kerberos
Authentication protocol for cluster identities — in CDP/Hadoop often the bridge from corporate directory to service principals and Ranger.
Apache Knox
Perimeter gateway for Hadoop/CDP services — centralizes authn/authz at the cluster edge and decouples clients from internal endpoints.
Soft Delete
Mark a record deleted instead of physically removing it — helps audit and CDC, complicates queries.
Cloudera CDP
Cloudera Data Platform for hybrid/multi-cloud and on-prem analytics — SDX bundles Shared Data Experience (Ranger, Atlas, HMS) as a governance surface, but does not replace a business operating model.
Time Travel
Querying a prior table version (lakehouse/warehouse) — useful for audit and rollback, costs retention.
Cloudera SDX
Shared Data Experience in CDP — shared security, governance and metadata services (including Ranger, Atlas, Metastore) across workloads.
Zero-Copy Clone
Clone of a table/DB without duplicating data — ideal for dev/test when lifecycle and rights are clear.
Apache Impala
Massively parallel SQL engine on Hive/HDFS/Ozone data — often the interactive analytics path in CDP alongside Spark and Hive.
Partition Pruning
Query reads only relevant partitions — requires clean partition keys and filters in contracts.
Apache ZooKeeper
Coordination service for distributed systems — quorums, leader election and config for Kafka, HDFS NameNode HA and other CDP services.
Star Schema
Fact table surrounded by denormalized dimensions.
Cloudera Manager
Operations console for Cloudera clusters — service lifecycle, config and health; a custodian tool, not business authority.
Predicate Pushdown
Filters execute as close to storage/engine as possible — less I/O, faster marts and BI.
Apache HBase
Distributed wide-column store on HDFS — its own namespace/table resources for Ranger and often tied to operational workloads.
Compaction
Merging small lake files and cleanup — keeps scan cost and time travel in check.
Apache Hudi
Lakehouse table format with upserts/incremental processing — alternative/complement to Delta and Iceberg.
LDAP / Corporate Directory
Enterprise directory for identities and groups — mapping to Ranger policies and service accounts is the critical governance link, not the group name alone.
Avro
Row-oriented serialization format with schema — common in streaming and schema registries.
Schema Evolution
Planned schema change without uncontrolled breaks — needs compatibility rules and contracts.
Workflow Orchestrator
System for scheduling, dependencies, and retries of data jobs — not the same as transformation (dbt).
Ingestion Tool
Connector/ELT tool loading from SaaS/DBs into landing — replaces neither modeling nor contracts.
Protocol Buffers
Compact binary schema format — common in APIs and event streams with strict evolution.
ORC
Columnar format for analytical workloads — Parquet alternative, strong in Hive ecosystems.
Fact Table
Measurable events or transactions at a declared grain.
Apache Arrow
In-memory columnar format for fast exchange between engines — bridge across Spark, DuckDB, BI.
DuckDB
Embedded analytical DB for local/edge SQL workloads — great for exploration, not automatically an enterprise hub.
Log-Based CDC
Change data capture from database logs — low source impact, needs privileges and lag monitoring.
Exactly-Once
Processing promise with no duplicates and no loss — expensive; often effective exactly-once via idempotency is enough.
At-Least-Once
Every event arrives at least once — duplicates possible; downstream must be idempotent.
Replay
Reprocessing historical events/loads — needs idempotency and clear watermarks.
Data Share
Governed exchange of datasets across boundaries — without uncontrolled exports.
Clean Room
Controlled environment for joint analysis without raw data exchange — privacy by design.
Blue/Green Deployment
Two parallel environments; switch only after validation — related to parallel run/cutover.
Dimension Table
Descriptive context (customer, product, calendar] joined to facts.
Canary Release
New release first for a small traffic/user share — limit risk before full cutover.
Feature Flag
Runtime switch for features/pipelines — decouples deploy from release, needs governance.
Outbox Pattern
Publish events reliably from the same transaction as the state change — against message loss.
Infrastructure as Code
Infra versioned and reviewable in the repo — foundation for reproducible data platforms.
GitOps
Git as source of truth for deploy state — pull-based, auditable, good for platform standards.
CQRS
Separation of write and read models — useful under different write/read loads, increases complexity.
Saga
Sequence of local transactions with compensation instead of distributed 2PC — for multi-service flows.
Polyglot Persistence
Multiple storage technologies per use case — needs clear ownership and integration contracts.
Change Data Feed
Table change stream from lakehouse formats — downstream reads inserts/updates/deletes incrementally.
Conformed Dimension
Shared dimension meaning and keys reusable across marts and domains.
Z-Order
File clustering by multiple columns — improves skip/pruning for mixed filters.
Liquid Clustering
Flexible clustering strategy in modern lakehouses — less rigid partitioning.
Vacuum
Cleanup of unreferenced files after time-travel retention — cost and compliance lever.
Shallow Clone
Clone that shares metadata and does not copy data — fast, but lifecycle-coupled.
Deep Clone
Clone with an independent data copy — costlier, but decoupled from source lifecycle.
Data Marketplace
Catalog/marketplace to find and consume data products — needs certification and contracts.
SCD (Slowly Changing Dimension]
Pattern family for how dimension attributes change over time.
Zero-ETL
Architecture promise: analytics close to the source without classic batch ETL copies — governance and latency still matter.
Data Virtualization
Queries across distributed sources without physical consolidation — fast to integrate, often costly for performance and lineage.
Lambda Architecture
Parallel batch and speed layers for correctness plus low latency — duplicated logic is the classic cost.
Kappa Architecture
One stream pipeline for realtime and replay instead of separate batch/speed layers — simplifies code, needs a robust event log.
Event Sourcing
State is derived from an immutable event sequence — strong for audit, but replay and schema evolution are demanding.
SCD Type 2
Historize attribute changes with effective dating or version rows.
Message Broker
Middleware for async messages (topics/queues) — decouples producers and consumers in event architectures.
Event Choreography
Services react to events without central orchestrator — flexible but harder to debug than orchestration.
Event Orchestration
Central coordinator drives steps and retries across services — easier to monitor than pure choreography.
Saga Orchestrator
Component running distributed transactions as sagas with compensating actions — common in microservice stacks.
Warm Standby
Standby system is prepared and quick to start — balance between cost and recovery time vs. cold standby.
Active-Active
Multiple sites/clusters take traffic in parallel — highest availability, complex for data consistency.
Geo-Replication
Replicates data across regions for DR and latency — cost and consistency model must align.
Multi-Region
Architecture across multiple cloud regions — affects data residency, failover, and compliance together.
CDC (Change Data Capture]
Capture source inserts, updates, and deletes for incremental loads.
Delta / Incremental Load
Process only changes since last successful run (often via MERGE].
Surrogate Key
System-generated durable key independent of source natural keys.
Natural Key
Business or source identifier used for matching and lineage back to origin.
Golden Record
Resolved, governed master representation of an entity (e.g. Golden Customer].
As-Was History
Point-in-time view of relationships and attributes as they were at a past date.
Outbox Pattern
Writes domain event and DB change in one transaction; a relayer publishes async — avoids lost updates between DB and bus.
Saga Pattern
Orchestrates distributed steps with compensations instead of one global DB transaction — failure is part of the design.
Event-Driven Architecture
Systems react to events instead of sync call chains — decouples producers/consumers, needs clear event contracts.
Micro-Batch
Processes events in short intervals (seconds/minutes) instead of one-by-one — trade-off between latency and throughput.
Idempotency
Applying the same request/event multiple times yields the same end state — required with retries and exactly-/at-least-once.
Write-Audit-Publish
Write and audit data first, only then publish to consumers — prevents “broken goes live”.
CQRS Read Model
Read-optimized model separate from the write model — often denormalized and updated via events.
Data Replication
Copies data between systems for latency, resilience, or isolation — needs conflict and lag strategy.
Fan-Out
One event/request is delivered to many consumers — scales read load, increases observability need.
Append-Only
New entries are only appended, never overwritten — audit and replay get easier; corrections are new events.
Event-Carried State Transfer
Event carries enough state for consumers to work without querying the producer — trade-off: larger payloads.
Schema Drift
Unexpected source or schema change that breaks contracts or loads.
Strangler Fig Pattern
Replace legacy incrementally, route by route — avoid big-bang migration.
Bulkhead Pattern
Isolate resources so one failure does not take everything down — e.g. separate pools per tenant/domain.
Referential Integrity
Guarantee that foreign keys point to existing parents — central for facts, dimensions, and DQ gates.
Sidecar Pattern
Helper process runs beside the main service (logging, mesh, auth) — decouples cross-cutting from app code.
Bridge Table
Helper table for many-to-many between facts and dimensions (or among dimensions) without breaking grain.
Service Mesh
Infrastructure layer for service-to-service (mTLS, routing, observability) — policy without app changes.
Materialized View
Stored query results for fast reads — refresh strategy (incremental/full) is part of the design.
Read Replica
Readable copy of the primary DB — offloads analytics; replication lag must be communicated.
Snowflake Schema
Dimensional variant with normalized dimension tables — saves redundancy, often costs join complexity vs. star.
Cold Storage
Cheap, slow storage tier for rarely used data — retrieval time belongs in the SLA.
Anti-Corruption Layer
Translates foreign/legacy model into your domain model — stops legacy language infecting the design.
Junk Dimension
Bundles low-cardinality flags/codes into one dimension — keeps the fact table lean.
Degenerate Dimension
Dimensional attribute that lives only on the fact table (e.g. invoice number) — without its own dimension table.
Hexagonal Architecture
Domain core in the center, I/O via ports/adapters — testable and swappable without bending the core.
DAX
Formula and query language for Power BI / Analysis Services — measures and calculated columns in the semantic model.
Ports and Adapters
Port defines interface, adapter implements it (DB, API, UI) — core knows no SQL details.
Event Backbone
Central event infrastructure domains plug into — scaling and schema governance are shared duties.
SCD Type 1
Dimension change overwrites the current value — no history; simple and often wrong for audit.
Log Shipping
Ships transaction logs to standby/replica — classic for DR; latency and failover must be planned.
SCD Type 3
Stores limited prior versions in extra columns — rare; usually SCD2 or as-was is better.
Hot Path
Path for latency-critical processing (real-time/near-RT) — costlier, must stay lean.
Role-Playing Dimension
Same dimension used multiple times in different roles (order date, ship date) — without duplicating physical tables.
Cold Path
Path for batch/archive processing with higher latency — cheaper, for history and heavy aggregation.
Set Analysis
Qlik syntax to compute aggregations independently of — or deliberately against — the current selection state.
Backpressure
Downstream signals “slow down” so upstream does not overflow — without backpressure, streaming crashes.
Factless Fact Table
Fact table without numeric measures — captures events or coverage (e.g. attendance, promotion coverage).
Late-Arriving Data
Facts or dimensions arrive after the expected load window — incremental/SCD strategies must tolerate it.
Matching
Decision of which source records are the same entity — prerequisite for golden record and MDM.
LookML
Looker’s modeling language for views, explores, and measures — business logic as code instead of inside every report.
Historized Entity
Entity with a validity timeline (valid from/to) — core of the integrated core and as-was reporting.
Bus Matrix
Matrix of business processes × conformed dimensions — planning tool for marts and shared dimensions.
Additive Measure
Measure that can be summed across all dimensions (e.g. revenue) — foundation of clean aggregations.
LOD Expression
Tableau expressions that control aggregation level independently of the current viz grain.
Semi-Additive Measure
Summable across some dimensions, often not across time (balances) — needs snapshot logic.
Non-Additive Measure
Not meaningfully summable (ratios, distinct counts) — aggregation must be recalculated.
Transaction Fact
Fact per business event (order line) — fine grain, additive measures.
Filter Context
The set of active filters/selections under which a measure formula evaluates — the source of many “wrong” KPI numbers.
Periodic Snapshot Fact
Fact with a regular status picture (day/month) — typical for balances and semi-additive measures.
Accumulating Snapshot Fact
Fact that progresses a process across milestones (apply → approve → close) — rows are updated.
Kimball
Dimensional approach with bus matrix and conformed dimensions — bottom-up marts around business process.
Master Measure
Reusable, versioned measure definition in the semantic layer — not reinvented per report.
Inmon
Top-down enterprise DW with normalized core and dependent marts — contrast to the Kimball bus.
Canonical Model
Shared integration model across systems — reduces point-to-point, needs strong ownership.
Many-to-Many
Relationship where both sides have multiple partners — in dimensional models often via bridge tables.
Calculation Group
Tabular/Power BI pattern to model time intelligence or presentation variants centrally instead of measure explosion.
Cardinality
How many partners a relationship has (1:1, 1:n, n:m) — drives joins, grain, and BI filter paths.
Sparsity
Many combinations without events — affects storage, aggregations, and BI performance.
Outrigger Dimension
Dimension hanging off another dimension instead of the fact — use sparingly.
Self-Service BI
Business users build reports themselves — works only with governed semantics and clear consumption contracts, otherwise shadow IT.
Mini-Dimension
Small dimension for rapidly changing attributes — relieves large SCD2 dimensions.
Late-Arriving Dimension
Fact arrives before the dimension — needs inferred/unknown members and later matching.
Multi-Valued Dimension
One fact has multiple dimension values at once — typically solved via bridge/weight factors.
Embedded Analytics
Analytics and visuals embedded in operational apps — same contracts and entitlements as standalone BI.
SCD Type 6
Hybrid of types 1/2/3 — current and historical views in parallel, modeling-heavy.
SCD Type 0
Attribute stays unchanged (retain original) — rare; document clearly.
Heterogeneous Dimension
One dimension table for different entity types with shared keys — use carefully.
Metric Lineage
Traceability from measure/KPI back to sources, transformations, and owners — mandatory for trusted metrics.
Certified Dataset
Approved, reviewed semantic dataset for self-service — a badge does not replace owner and contract.
Slowly Changing Dimension
Pattern for dimension changes over time (Type 1 overwrites, Type 2 historizes) — without a strategy every trend analysis lies.
DirectQuery
BI query mode against the source instead of import — choose freshness vs. performance and governance trade-offs deliberately.
Thin Consumer Interface
Reports hold little business logic of their own — they consume measures and contracts from the semantic layer.
Headless BI
Semantics and metrics as an API/service — visualization is swappable, definitions stay central.
OLAP
Analytical multidimensionality (slice/dice, hierarchies) — today often as tabular/columnar engines instead of classic cubes.
Tabular Model
Columnar semantic model (e.g. Power BI / AAS) with relationships, measures, and often import/DirectQuery.
Shadow IT
Ungoverned reports, exports, and pipelines outside the platform — symptom of missing semantics and delivery.
Time Intelligence
Period comparisons and YTD/MTD logic — central in the semantic layer, not copy-pasted into every report.
Data Quality
Measurable fitness of data for a purpose — rules, ownership, monitoring, and remediation instead of one-off checks.
Composite Model
Semantic model with mixed storage modes — flexibility with higher governance and performance complexity.
Master Dimension
Reusable dimension/hierarchy in the semantic layer — analogous to a master measure for attributes.
Metric Layers
Separation of raw measure, business measure, and presentation measure — prevents logic in every dashboard.
Governed Self-Service
Self-service on certified datasets and clear contracts — freedom with guardrails instead of shadow IT.
Vanity Metric
Metric that looks good but changes no decision — symptom of missing KPI contracts.
North Star Metric
One central outcome metric for product/org — needs grain, owner, and supporting metrics.
Leading Indicator
Metric that predicts outcomes (pipeline, activation) — complements lagging indicators.
Lagging Indicator
Metric that measures outcomes after the fact (revenue, churn) — important, but alone too late to steer.
Metric Proliferation
Uncontrolled growth of similar measures — trust drops; certification and deprecation help.
Data Contract
Agreement between producer and consumer on schema, semantics, SLAs, and breaking-change rules.
Dashboard Sprawl
Too many similar dashboards without owner and dataset contract — classic symptom of self-service without guardrails.
Drill-Down
Navigate from coarse to fine hierarchy level — needs conformed hierarchies and clear grain.
Drill-Through
Jump from aggregate to detail rows or another report — lineage and entitlements must travel with it.
Ambiguous Path
Multiple filter paths between tables in the semantic model — yields wrong aggregates if unresolved.
Bidirectional Filter
Filters flow both ways on a relationship — powerful and risky for ambiguous paths.
Operational Reporting
Day-to-day reporting close to the process — often different latency and grain than analytical marts.
Analytical Reporting
Analysis over integrated, historized models — typically mart/semantic layer rather than raw ODS.
Dashboard
Consumption surface with visuals, filters, and context for a concrete decision or monitoring task. A dashboard should consume approved metrics; it is not a data catalog, metadata source, or suitable home for competing core logic.
Scorecard
Compact KPI overview with targets/status — needs contracts, or you get colorful lights without trust.
Fitness for Purpose
Quality judged against an explicit use case — not absolute perfection.
OKR
Goal system of objectives and measurable key results — key results are often KPIs with owner and grain.
Cohort
Group sharing a start trait (signup month) — analyses need stable grain and history.
Attribution
Assigning outcomes to touchpoints/causes — highly sensitive to model and definition.
Workspace
Collaboration and publish boundary in BI platforms — entitlements, promotion, and certification hang off it.
Dataset
Reusable semantic model/package for reports — ideal point for certification and contracts.
Promotion
Moving content/datasets from dev/test to prod — needs checks, owner, and rollback.
Excel Export
Export from governed surfaces into spreadsheets — often the start of shadow IT and the two-version problem.
Manual Reconciliation
Manually reconciling numbers across systems — symptom of missing contracts, grain, and lineage.
Stale Metric
Metric with outdated definition or no owner — trust killer despite looking “official”.
Data Observability
Detect unexpected volume, null, distribution, schema, and freshness anomalies beyond fixed rules.
Conflicting Metric
Two measures answer the same question differently — the two-version problem in numbers.
Orphan Report
Report without owner or usage — clean up, archive, or deprecate.
Unused Dataset
Dataset without consumers — cost and risk without value; usage metadata makes it visible.
Spreadsheet Hell
Critical logic and truths only in spreadsheets — end state of shadow IT and Excel export.
DQ Gate
Hard stop or promote condition in the pipeline based on quality outcomes.
Field Parameter
BI control that lets users switch fields/metrics in visuals — governance must constrain allowed dimensions.
Dual Storage Mode
A table can use import and DirectQuery together — flexible, but harder to explain and debug.
Incremental Refresh
Loads only new/changed partitions instead of a full reload — saves time, needs stable keys and archive policy.
Live Connection
The report hits the semantic model live — always current data, but performance and model governance sit centrally.
Import Mode
Data is loaded into the BI model and cached there — fast for users, refresh windows and memory become critical.
Rule Registry
Catalog of DQ checks, owners, severity, and execution context.
Remediation
Owned fix-and-validate loop for quality and metadata defects.
Staging Area
Intermediate zone for raw or lightly transformed data before core model — separates ingest from business logic.
Persistent Staging
Staging kept with history — enables rebuilds and audit without re-pulling sources.
Incremental Pipeline
Processes only new/changed data since last run — standard for scale, needs clean keys and watermarks.
Ragged Hierarchy
Hierarchy with varying depths (levels skipped) — e.g. region without intermediate country node.
Unbalanced Hierarchy
Nodes have different child depths — reporting and drill must model parent-child paths explicitly.
Parent-Child Hierarchy
Hierarchy via parent ID per row — flexible but queries and BI drill often costlier than snowflake dims.
Hierarchy Path
Stored path from root to node — speeds drill and filter in unbalanced hierarchies.
Freshness
How current data (or metadata] is versus the agreed expectation.
Completeness
Required fields and records present for the declared purpose.
Consistency
Same meaning and rules across systems, layers, and replicas.
Accuracy
Values correctly represent the real-world fact for the use case.
Uniqueness
No unwanted duplicates at the declared grain or key.
Snapshot Fact
Fact table measures state at fixed points in time (e.g. daily inventory) — not every individual transaction.
Validity
Values conform to allowed formats, domains, and reference sets.
Accumulating Snapshot
Fact row follows a process across milestones (order→ship→pay) and is updated — typical for cycle times.
SCD Type 1
Dimension attribute is overwritten — history is lost, suitable for corrections with no analysis need.
Late-Arriving Fact
Fact appears after the expected load time — pipelines must handle backfill/replay and dimension keys cleanly.
SCD Type 3
Keeps current and previous attribute values in columns — limited history, not full SCD2.
Galaxy Schema
Multiple fact tables share dimensions — common in large DWHs, needs conformed dimensions.
Rapidly Changing Dimension
Dimension attribute changes so often that classic SCD2 explodes — often mini-dimension or outrigger instead of full history.
Hard Delete
Record is physically removed — irreversible; often needed for erasure, hard on history/audit.
Additive Measure
Metric may be summed across dimensions (revenue, quantity) — averages are often not additive.
Semi-Additive Measure
Summable across some dimensions, not time (balance, inventory) — grain decides.
Non-Additive Measure
Must not be naively summed across rows (ratio, avg price) — aggregation needs a formula, not SUM.
Root Cause Analysis (DQ]
Trace defect to originating system or process instead of patching marts.
Tombstone Record
Marker for deleted entries in logs/streams — consumers must propagate deletes, not ignore them.
Bronze Layer
Raw/landing zone with minimal transformation — audit and replay, no business logic.
Silver Layer
Cleansed, conformed data — joins and quality, not yet mart-specific.
Gold Layer
Business-ready marts/metrics for consumers — grain and KPI definition live here.
Measure Group
Group of related measures/facts sharing grain — wrong grouping mixes KPIs.
Domain-Driven Design
Model around the business domain (bounded contexts, ubiquitous language) — not the other way around from tools.
Quality by Design
Bake quality into contracts, models, and pipelines — not “find” it later by sampling dashboards.
Anomaly Detection
Automatic detection of unexpected patterns in volume, distributions, or metric values — complements rule-based tests.
Data Profiling
Statistical inventory of columns (nulls, cardinality, patterns) — basis for rules and contracts.
Contract as Code
Data contracts and assertions versioned in the repo and checked in CI — not only as a wiki paragraph.
Timeliness
DQ dimension: data arrives in time for the decision — related to freshness, but judged against the use case.
Incident Runbook
Step-by-step guide for pipeline/DQ incidents — owners, checks, escalation, communication.
DataOps
DevOps practices for data products: CI/CD, observability, short feedback loops between produce and consume.
Quality Score
Aggregated score from DQ rules/dimensions — useful as a trend, dangerous as the only truth.
Assertion
Machine-checkable expectation on data (row counts, ranges, references) — building block of contract-as-code.
On-Call
Rotating responsibility for off-hours incidents — needs runbooks and clear escalation.
MTTR
Average time to recover after incidents — runbooks and on-call reduce it.
Error Budget
Allowed unreliability under the SLO — steers change pace vs. stability.
Toil
Manual, repetitive ops work without lasting value — automation and golden paths reduce it.
PII
Personally identifiable or linkable data. Needs classification, masking, purpose binding, and proven deletion/restriction paths.
Orphan Pipeline
Pipeline without clear owner/consumer — costs money and creates silent incidents.
DSDR
Process and technical ability to execute and evidence data-subject rights (erase/restrict) across systems and lineage.
Quarantine Zone
Isolated holding area for failed records — pipeline continues, bad data does not reach consumers.
Reconciliation Check
Compares sums/counts between source and target — finds silent loss that row tests miss.
Data Reliability
Overall sense that data arrives on time, complete, and trustworthy — more than single DQ checks.
Volume Anomaly
Unexpected jump or drop in row count — often the first signal of broken upstream jobs.
Retention
Rules for how long data may stay active or archived — separated from backup and analytical marts.
Masking
Technique to hide or replace sensitive values for unauthorized roles — ideally policy-driven and lineage-aware.
Data Classification
Labeling sensitivity or purpose class that drives protection and use.
Sensitivity Label
Platform label (e.g. Purview] binding policy to assets or columns.
Purpose Limitation
Use only for the agreed purpose — before tooling and mart design.
Report Book
Bundled report collection with shared navigation — typical for PDF/print compliance packs.
Mobile BI
BI for phone/tablet — layout, offline, and RLS must be designed for small screens.
Embedded Report
Report/dashboard embedded in product UI — needs embedding API, auth, and consistent semantic layer.
Report Subscription Email
Automated report delivery via email on schedule — distribution layer, not a substitute for interactive BI.
Scheduled Refresh
Schedule for data refresh in BI tool — SLA for “fresh” dashboards, often tied to pipeline SLA.
Tokenization
Replace sensitive values with reversible tokens under controlled vaulting.
Pseudonymization
Reduce identifiability while allowing controlled re-link under safeguards.
Anonymization
Irreversible removal of personal identifiability for a stated threat model.
Redaction
Drop or blank high-risk fields so they never reach curated or mart layers.
Workforce / Employee Data Policy
Separate handling rules for workforce identity vs customer PII in RAW→Mart.
GDPR
EU privacy framework with purpose limitation, data-subject rights, and accountability — drives classification, retention, and DSDR.
Legal Hold
Freeze against deletion/change due to litigation or investigation — overrides normal retention.
Dynamic Masking
Masking at query time by role — raw data stays stored, visibility is controlled.
Hashing
One-way transform of identifiers — often for joins without plaintext PII, with collision and rainbow risks.
Aggregation Table
Precomputed summary at a coarser grain — speeds reports, must stay consistent with detail source.
Data Minimization
Store/share only necessary attributes and rows — tightly linked to purpose limitation and least privilege.
Consent
Freely given, informed permission to process — one lawful basis among others, not the only one.
Direct Lake
BI reads Parquet/Delta in the lake without classic import — fresh like DirectQuery, often faster than pure warehouse import.
Lawful Basis
Legal ground for processing (consent, contract, legitimate interest…) — must fit the purpose — masking is a control, not a lawful basis.
Object-Level Security
Hides tables/columns/measures from roles — complements RLS (rows), does not replace it.
Controller
Party that determines purposes and means of processing — carries accountability to data subjects.
Drillthrough
Navigation from aggregated view to a detail page with context filters — not a substitute for clean grain definition.
Processor
Processes personal data on behalf of the controller — needs a contract and evidenced controls.
Tooltip Page
Small report page as hover detail — gives context without page navigation, should stay lean.
Special Category Data
Especially protected data (health, biometrics…) — stricter requirements than ordinary PII.
What-If Parameter
Interactive parameter for scenarios (e.g. price ±10%) — simulates, does not persist facts in the source.
Calculated Column
Column materialized in the model (row context) — unlike a measure (filter context); memory-heavy on large tables.
Synthetic Data
Artificially generated data for test/training — reduces PII risk, needs realism checks.
Differential Privacy
Mathematical privacy protection via controlled noise — strong, but utility trade-off.
Visual-Level Filter
Filter applies to one visual only — not the whole page/report; easy to miss when debugging.
Data Residency
Requirement for where data may physically/legally reside — drives cloud region and sharing design.
Page-Level Filter
Filter for all visuals on a report page — scope between visual- and report-level.
PII Scan
Automated search for personal patterns in schemas/content — starting point for classification.
Report-Level Filter
Filter applies report-wide across pages — powerful, but risky when users do not see the scope.
Sync Slicer
Slicer selection stays synced across pages — less clicking, more coupling.
Tombstone
Marker that a record is deleted/suppressed — relevant for CDC, DSDR, and soft deletes.
k-Anonymity
Each quasi-identifier profile appears at least k times — classic, limited anonymization approach.
Paginated Report
Print/PDF-oriented report with fixed layout — unlike interactive dashboards.
Format-Preserving Encryption
Encryption that preserves format/length — for legacy fields; does not replace access policies.
On-Premises Data Gateway
Bridge from cloud BI to on-prem sources — needs HA, permissions, and monitoring like production infra.
Bookmark State
Saved filters/views in reports — personal vs shared must be clear or “wrong truth” spreads.
Homomorphic Encryption
Compute on encrypted data without decrypting — powerful, still often costly/complex today.
Encryption at Rest
Encryption of stored data — baseline hygiene; does not protect against authorized queries.
Report Theme
Central colors/fonts for reports — without a theme, every dashboard becomes a design experiment.
Encryption in Transit
Encryption on the wire — standard for all pipeline and API paths.
Mobile Report
Report layout for small screens — different priorities than desktop, not just scaled dashboard.
Health Data
Health-related personal data — usually special category / especially protected.
Report Subscription
Automated delivery of report snapshots — snapshot time and filters must be documented.
Employee Data
Personal employee data — own policies, co-determination, and purpose limits.
Visual Cross-Filter
Click in one visual filters others — powerful for exploration, easy to miss when debugging.
Button Slicer
Filter as clickable buttons/tiles — UX-friendly, but limited cardinality recommended.
Relative Date Filter
Filter like “last 7 days” relative to today — dynamic; timezone must be correct.
Data Clean Room
Controlled environment for joint analytics without raw data exchange — queries yes, copying datasets no.
Top N Filter
Shows only top/bottom N by measure — tie-breaker and filter context must be documented.
Retention Schedule
Defines how long each data class is kept and when it is deleted/archived — without a plan compliance risk grows.
Small Multiples
Same chart repeated per category — compares patterns, not absolute scale alone.
Bookmark Navigation
Buttons jump to saved filter/page states — guided analytics instead of free exploration.
Right to Erasure
Right to have personal data erased — needs lineage through backups and downstream copies.
RBAC
Access by role membership.
ABAC
Access by attributes (clearance, purpose, residency, etc.].
Least Privilege
Minimum access needed for the job — default deny elsewhere.
Segregation of Duties (SoD]
Split conflicting duties (e.g. grant vs approve] to reduce abuse risk.
Section Access
Qlik row-reduction security model in apps.
Row-Level Security (RLS]
Filter rows by user or role claims in warehouse or BI.
Access Recertification
Periodic re-approval that entitlements are still needed.
Distribution Check
Checks whether value/category distribution looks expected — early drift alarm before hard business-rule fails.
Uniqueness Constraint
Rule that keys or combinations must be unique — basis for referential integrity and aggregation.
Referential Check
Validates foreign keys point to existing parent rows — prevents orphaned facts in star schemas.
Schema Validation
Checks columns, types, and required fields against expected schema — often gate before lakehouse write.
Anomaly Detection Rule
Automated rule for unusual values or counts — complements static thresholds on volatile metrics.
Profiling Job
Scheduled job for null rates, distinct counts, and patterns — input for quality rules and anomaly detection.
Quality Gate
Checkpoint in CI/CD/pipeline — deploy or promote only when defined DQ checks pass.
Observability Alert
Alert from metrics/logs (latency, freshness, error rate) — connects ops visibility with data quality.
IAM
Identity and access management backbone for authenticating subjects to data systems.
Column-Level Security
Access control at column grain — complements RLS when entire attributes must be invisible to roles.
Encryption
Protection of data in transit and at rest — baseline hygiene; replaces neither masking nor RBAC.
Entitlement
Concrete access right of an identity to an asset — subject of recertification and least privilege.
Service Principal
Non-human identity for pipelines and apps — own entitlements, rotation, and recertification.
Audit Log
Evidence of who accessed or changed what when — basis for recertification and incidents.
SSO
One login for many apps — simplifies IAM, but also centralizes risk.
MFA
Multiple authentication factors — baseline for privileged and PII access.
SCIM
Standard for automated provisioning/deprovisioning of identities into apps.
Zero Trust
No implicit trust from network zone — continuous verification of identity and context.
Secrets Management
Central, rotatable store for keys/tokens — no secrets in repos or notebooks.
Key Vault
Managed service for keys and secrets — often the anchor for encryption and BYOK.
BYOK
Customer controls encryption keys — increases control and compliance requirements.
Row Access Policy
Policy object controlling row access in the warehouse — complements app-side RLS.
Object Tag
Metadata tag on tables/columns for classification and policy binding — active metadata in action.
Just-in-Time Access
Time-bound rights only when needed — reduces standing privileges.
PAM
Management of highly privileged access — vaulting, session control, recording.
Break-Glass Access
Emergency access with strong audit — exception, not a steady state.
Column Masking Policy
Policy that role-masks column values — often bound to tags/classification.
Tag-Based Access
Access via classification/object tags instead of object names alone — scales better in large estates.
Lineage
Traceable origin and transformation of data — from source through pipelines to reports and deletion paths.
Data Catalog
Search and relationship surface for data assets, terms, owners, policies, and consumers. A data catalog makes metadata discoverable and connected; it is not automatically the system of record for every metadata item and not a dashboard or semantic layer.
Freshness SLA
Promised maximum data age until use — without measurement, “fresh” is marketing.
Metadata
Descriptive and controlling facts about data, reports, models, or processes: schema, meaning, owners, classification, lineage, quality, policy status, and time context. Metadata is the governance content; the catalog is only one possible surface for it.
Null Rate
Share of missing values in a column/time window — early warning for broken feeds and schema drift.
Data Diff
Compares two datasets/versions at row or aggregate level — finds drift after refactors and migrations.
Expectation Suite
Versioned set of declarative quality rules (range, uniqueness, references) — tests as code, not ticket prose.
Completeness Check
Checks whether expected fields/rows are present — “correct but incomplete” still fails.
Metadata Catalog
Catalog view that indexes metadata from different sources and makes it usable. The term usually describes the surface or platform, not automatically the business authority for every definition, classification, or approval.
Uniqueness Check
Ensures keys/combinations are unique — duplicates break joins and KPIs.
Control Plane
Control layer that manages rules, policies, identities, approvals, lineage, or operating states. It may prepare decisions or enforce them technically, but it does not replace the accountability of owner, steward, or custodian.
Validity Rule
Value must fall in an allowed domain (range, enum, format) — “not null” alone is not enough.
Duplicate Rate
Share of duplicate keys/rows in a window — early signal for broken upserts and CDC gaps.
Orphan Record
Row without a valid parent/dimension reference — joins yield gaps or wrong “unknown” buckets.
Accuracy Check
Checks whether values are factually correct (reconcile with source/golden set) — technically “green” is not enough.
Business Glossary
Agreed business terms and definitions — related to, not identical with, the catalog.
Consistency Check
Same metric/entity aligns across systems — DWH vs CRM contradictions are classic.
Conformity Check
Data matches format/standard (ISO date, currency code) — schema ok, content still unusable.
Integrity Check
Relationships between entities hold (FK, cardinality) — orphan rows are an integrity fail.
Shift-Left Quality
Quality rules and tests as early as pipeline/dev — avoid expensive fixes in the dashboard.
Timeliness Check
Checks whether data arrives on time — distinct from freshness SLA (age since last load).
Volume Check
Monitors expected row volume per load — sudden halving is often worse than a schema error.
Range Check
Values must fall in allowed interval — catches unit errors and overflow early.
Data Quality Scorecard
Aggregated view of DQ dimensions per product — score without action plan is decoration.
Quality Dimension
Axis of data quality (accuracy, completeness, timeliness …) — shared vocabulary for scorecards.
Active Metadata
Metadata that drives automation (policies, quality, routing], not only documentation.
Metadata Provenance
Who or what authored a metadata claim and from which source of truth.
Metadata Harvesting
Automated collection of technical metadata from platforms into the catalog.
Impact Analysis
Trace downstream blast radius of a field, model, or KPI change via lineage.
Column Lineage
Field-level origin and transform path (needed for PII propagation and DSDR].
Metadata Enrichment
Add business context, owners, classification, and KPI links to technical assets.
AI-Ready Metadata
Complete, current, permitted-use metadata suitable for assistants and RAG.
dbt meta
Structured metadata in dbt YAML used to drive governance automation.
Centralized vs Federated Metadata
Decide per capability what is central discovery vs domain-authored truth.
Privacy by Design
Build privacy into architecture and process from the start — not as a last filter before go-live.
Data Masking
Obscures sensitive values in non-prod/self-service (hash, partial, fake) — not a substitute for prod access control.
DPIA
Structured risk analysis before high-risk processing — documents mitigations, does not replace ongoing controls.
Consent Management
Captures, versions, and propagates consent in machine-readable form — a marketing opt-in alone is not proof.
Controller vs Processor
Controller decides purposes/means; processor acts on instructions — mixing them up costs contracts and liability.
Legitimate Interest
Legal basis after balancing interest vs data-subject rights — not a free pass without documentation.
OpenLineage
Open standard for job and dataset lineage events — interoperable across orchestrators and catalogs.
Data Subject
Natural person to whom personal data relate — holder of access and erasure rights.
Cross-Border Transfer
Transfer of personal data to other jurisdictions — needs a transfer mechanism and risk analysis.
Schrems II
CJEU ruling on US transfers: Privacy Shield invalid — transfer impact assessments and supplementary measures became central.
Retention Policy
Rules for how long which data is kept for what purpose — without technical enforcement, policy is folklore.
Data Breach
Unauthorized disclosure/access to personal data — notification deadlines and forensics are mandatory.
Technical Metadata
Schema, types, jobs, storage settings — what systems know about data structures and pipelines.
Breach Notification
Duty to inform authority/subjects about relevant incidents — needs playbook and evidence.
Re-Identification Risk
Risk of linking pseudonymized/anonymized data back to a person — watch quasi-identifiers.
Standard Contractual Clauses
Contract modules for third-country transfers — complement TIA and technical measures, do not replace them.
Data Processing Agreement
Contract between controller and processor on purpose, subprocessors, deletion — cloud without DPA is a red flag.
Privacy by Default
Most restrictive privacy setting is default — users must actively expand, not fight opt-out.
Transfer Impact Assessment
Risk analysis for third-country transfers post-Schrems II — complements SCCs, does not replace them.
Sub-Processor
Processor engages another vendor — chain must be transparent and controllable in the DPA.
Consent Record
Proof of who consented to what and when — must propagate to downstream systems.
Business Metadata
Business meaning, owners, glossary terms, and usage hints — makes technical assets understandable.
Operational Metadata
Runtimes, status, incidents, SLAs — metadata from operating pipelines and jobs.
Usage Metadata
Who uses which assets how often — basis for prioritization, certification, and cleanup.
Data Subject Request
Individual request for access, deletion, or correction — must be traceable across systems and lineage.
DSR Workflow
Standard process from ticket to deletion/export with deadlines — without workflow, copies linger in backups.
Opt-out
Objection to use (marketing, profiling) — must be reflected technically in segments and pipelines.
Opt-in
User actively consents — standard for many marketing and analytics use cases in the EU.
Legitimate Interest Assessment
Documented balancing for processing without consent — needs purpose, necessity, and balancing.
Declared Metadata
Explicitly documented or code-declared metadata (e.g. dbt meta, glossary) — intent, not observation.
Detected Metadata
Harvested or inferred metadata from systems — schema, lineage, usage — complements declared metadata.
Control-Driving Metadata
Metadata that actively drives policies, masking, routing, or gates — not merely describes.
RACI
Role matrix for decisions: who executes (R], who owns (A], who is consulted (C], who is informed (I].
Metadata Graph
Connected view of assets, owners, lineage, and policies — foundation for impact analysis and active metadata.
Business Lineage
Lineage in business terms (KPI ← mart ← domain) — understandable for owners and stewards, not only engineers.
Technical Lineage
Lineage at table/column/job grain from harvesting and runtime — precise for impact and debugging.
Enterprise Vocabulary
Shared business language across domains — mapping local labels to canonical glossary terms.
KPI Governance
Clear definition, owner, and change process for metrics — prevents conflicting numbers across tools and meetings.
Metadata Product
Metadata with product thinking: owner, SLA, versioning, and consumers — not just a catalog dump.
Schema Registry
Central management and evolution of event/message schemas — compatibility checks before breaking changes.
Event-Driven Metadata
Metadata updates as events from jobs/catalogs — instead of only nightly full harvests.
Static Metadata
Slow-changing metadata (owner, glossary, classification) — often declared, not harvested every minute.
Dynamic Metadata
Frequently changing metadata (freshness, usage, job status) — typically detected/observed from runtime.
Decision Rights
Who may decide purpose, access, definitions, and exceptions — and at what risk tier.
Privileged Access Management
Controls and monitors highly privileged accounts (admin, break-glass) — just-in-time instead of standing admin.
Service Account
Non-human identity for pipelines/apps — needs rotation, scope, and an owner, not a shared wiki password.
Source-Local Metadata
Metadata that originates and is maintained in the source tool — harvesting pulls it into the graph.
Break-Glass Access
Time-boxed emergency access with audit — an exception reviewed afterward, not an everyday path.
Descriptive Metadata
Metadata that explains and helps discovery — unlike control-driving metadata that steers systems.
Workload Identity
Identity for workloads (pods, jobs) without long-lived secrets — short-lived tokens instead of passwords in repos.
Conditional Access
Access depends on signals (device, location, risk) — MFA alone is often too coarse.
Distributed Metadata
Metadata across many tools/domains without a mandatory single catalog — needs federation and a clear source of truth per aspect.
Secrets Rotation
Regular replacement of keys/passwords — without automation, secrets stay “valid forever”.
Bounded Context
Bounded meaning space for terms and models — prevents “customer” meaning the same thing everywhere.
Identity Provider
Central authority for authentication/SSO — apps trust tokens instead of their own password DBs.
OAuth Client
Registered app with client ID/secret and scopes — rotation and least-privilege scopes are mandatory.
API Gateway
Central entry for APIs (auth, rate limit, routing) — policies here instead of duplicating in every service.
Mutual TLS
Client and server authenticate via certificates — stronger than TLS-only; needs rotation and CA governance.
Data Dictionary
Technical catalog of tables/columns/types — complements the business glossary, does not replace it.
Network Segmentation
Split network into zones to hinder lateral movement — keep data plane and admin separate.
Operating Model
Cadence, handoffs, capacity, and escalation that make roles real.
Passive Metadata
Automatically harvested schemas/stats without active curation — good for discovery, weak for binding semantics.
Security Posture
Overall picture of controls and gaps — score alone is not enough without prioritized remediation.
Threat Modeling
Systematic analysis of threats and mitigations — before go-live, not after the incident.
Attack Surface
Sum of exposed entry points (APIs, accounts, data exports) — shrinking beats more alerts.
Public Key Infrastructure
Infrastructure for certificates and trust chains — foundation for mTLS and signed tokens.
Phishing-Resistant MFA
MFA that resists replay/phishing (e.g. FIDO2) — SMS codes do not count.
Escalation Path
Defined route when Steward, Owner, or Platform cannot resolve within SLA.
Stewardship Intake
Prioritized entry path for definition, DQ, and classification work.
Stewardship Capacity Model
FTE or protected-time model so stewardship is funded, not „on the side“.
Governance Cadence
Recurring forums and reviews (classification, access, KPI, council].
Data Lifecycle
Create → use → retain/archive → delete/retire with accountable stages.
Role Sprawl
Too many overlapping RACI hats that dilute accountability.
Federated Governance
Central standards plus domain execution (vs pure central or pure local].
Policy as Code
Enforceable access, quality, and privacy rules in versioned, testable artifacts.
Data Technical Debt
Accumulated compromises in pipelines, models, and reports — interest paid as incidents, shadow IT, and slow change.
Accountability
Who finally owns outcome and risk — the “A” in RACI, not the same as doing the work (“R”).
KPI Operating Model
Cadence, roles, and change process around KPIs — from definition through certification to breaking changes.
ADR
Short record of an architecture decision including context and consequences — prevents tribal knowledge.
Tribal Knowledge
Critical knowledge only in a few heads — bus-factor risk; metadata and ADRs are antidotes.
GDR
Decision record of a governance decision: the decision, the A, scope and kill line — not meeting minutes.
Two-Version Problem
Two “official” numbers for the same question — typical without a metric store and decision rights.
SRE
Practice of running reliability with engineering — SLOs, error budgets, less toil.
Platform Engineering
Builds self-service platform products for teams — golden paths instead of ticket-driven ops.
Golden Path
Supported standard path (templates, CI, patterns) — fast and compliant at once.
Value Stream
End-to-end flow from need to outcome — helps prioritize data products and roadmaps.
Lead Time
Time from request to production-ready — core flow metric beside cycle time.
Cycle Time
Time of active work on an item — drops with smaller batches and less WIP.
WIP
Work open at the same time — high WIP lengthens lead time and creates half-done work without outcome.
Product Thinking
Focus on user outcomes, lifecycle, and ownership — instead of pure ticket/report production.
Community of Practice
Network for sharing standards and learning across teams — complements a CoE, does not replace it.
Bus Factor
How many people can be hit by a bus before knowledge/system fails — low with tribal knowledge.
Ticket-Driven Delivery
Work only via ticket queues — typical opposite of product thinking and golden paths.
Product-Driven Delivery
Delivery via products, outcomes, and roadmaps — instead of endless ad-hoc tickets.
Hero Culture
Systems run only via individual heroics — bus factor and burnout instead of reliability.
Hub and Spoke
Central hub plus domain spokes — org pattern for governance and platform.
Inner Source
Open-source practices internally — PRs, shared ownership, visible standards.
Trunk-Based Development
Short branches, frequent merges to trunk — fits CI/CD and small batches.
Canary Release
New version first for a small traffic share — limits blast radius; monitoring decides further rollout.
Shadow Deployment
New pipeline/model runs in parallel without user impact — compares results, switches later.
Chaos Engineering
Deliberately injecting failures in production/staging to prove resilience — not random breakage.
Circuit Breaker
Temporarily stops calls to an unhealthy downstream — prevents cascades, needs clear recovery rules.
Value Stream Mapping
Visualizes end-to-end flow from idea to value — makes queues and handoffs in data/BI visible.
Blue-Green Deployment
Two parallel environments; traffic switches atomically — faster rollback than patching live step by step.
Blameless Postmortem
After incidents, analyze causes and system gaps without blame — focus on fixes and learning.
Runbook
Concrete steps for ops/incident (checks, rollback, contacts) — a living document, not a one-off wiki page.
WIP Limit
Cap on parallel work — protects throughput and quality, including in data/BI backlogs.
Definition of Done
Shared done criteria (tests, docs, owner, monitoring) — “done” without a DoD is only locally done.
Definition of Ready
Criteria for when work is start-ready (scope, owner, data availability) — prevents half-baked starts.
Incident Severity
Incident severity rating (Sev1…n) drives response time and command structure — without it, everything is “urgent”.
Incident Commander
Coordinates response, communication, and decisions in an incident — not necessarily the deepest debugger.
Error Budget Policy
Defines what happens when the error budget is spent (freeze, reliability focus) — policy without consequence is theater.
Pair Programming
Two people on one problem (driver/navigator) — knowledge transfer and early review, not “twice the cost”.
Continuous Delivery
Keep software/data pipelines releasable at any time — manual “manual end-of-week deploy” is an anti-pattern.
Change Advisory Board
Forum for risky changes (prod, schema) — should reduce gatekeeping, not delay every ticket.
RTO
Maximum acceptable downtime until recovery — without RTO, DR is PowerPoint only.
RPO
Maximum tolerable data loss in time — backup frequency must match.
Ubiquitous Language
Shared, precise terms in team and code — prevents “revenue” synonyms in SQL and slides.
Context Map
Visualizes relationships between bounded contexts (upstream/downstream) — clarifies integration and ownership.
Trunk-Based Development
Short-lived branches, frequent merge to main — fits CI/CD and DataOps pipelines.
Game Day
Planned stress test/incident exercise in prod-like env — finds gaps before real outages.
On-Call Rotation
Schedule of who carries alerts/incidents when — without rotation, one hero burns out and knowledge silos.
Security Event
Logged security-relevant action (login, policy change) — raw material for SIEM and forensics.
SIEM Alert
Alert from SIEM correlation of multiple events — quality depends on parsing, tuning, and false-positive handling.
SOC Analyst
Triages alerts, escalates incidents, documents playbooks — operational heart of security operations.
Vulnerability Scan
Automated search for known weaknesses in images/hosts — findings need prioritization and patch SLA.
Penetration Test
Simulated attack on systems with report — deeper than scan but point-in-time, not continuous monitoring.
RAG (Retrieval-Augmented Generation]
Ground LLM answers in retrieved governed documents or data.
AI Guardrails
Controls that constrain prompts, tools, and outputs for safety and compliance.
Column Profiling
Statistics about columns (nulls, cardinality, patterns) — basis for DQ and catalog enrichment.
Prompt Injection
Attack that hijacks model behavior via malicious content in inputs or context.
Data Lineage Graph
Graph from sources→transforms→consumers — impact analysis needs edges, not wiki prose alone.
Open Metadata
Open platform for catalog, lineage, and governance workflows — metadata as a product, not a side-car.
Metadata Lineage
Metadata about origin and transformation — distinct from business-glossary prose alone.
Business Glossary Term
Business-defined term with owner and status — links to technical assets, does not replace them.
Semantic Metadata
Metadata about meaning (terms, KPIs, synonyms) — links glossary to technical assets.
Catalog Ingestion
Automatic loading of technical metadata into the catalog — without enrichment, just inventory.
Hallucination
Confident model output not grounded in retrieved or training evidence.
Human-in-the-Loop
Mandatory human review or approval for high-risk AI actions.
Training Data
Datasets used to train or fine-tune models — need lineage, rights, and quality.
Machine Unlearning
Methods to reduce or remove the influence of specific training points on a model — usually approximate in practice and paired with verify; often an alternative or complement to retrain/adapter drop and index rebuild.
AI Ingestion Gate
Control point before retrieval index or training ingest: hygiene disposition, PII/secret rate, purpose approval, and metadata contract must pass — otherwise block or a time-boxed waiver.
Shadow AI
AI use outside sanctioned paths — e.g. public chats, personal API keys, uncontrolled vendor POCs — without owner, purpose card, and evidence.
Sanctioned AI Path
Catalogued AI usage path with purpose, allowed data classes, owner/steward, risk class, logging, and exit criteria — prerequisite for production operation.
Corpus Poisoning
Intentional or negligent contamination of retrieval/training corpora (e.g. prompt-injected docs, tainted vendor packs) — needs trust tiers and integrity gates, not hygiene alone.
Inference
Runtime model execution against new inputs.
Feature Store
Governed reuse of ML features with versioning and serving contracts.
LLM
Large language model for text/code — needs guardrails, grounding, and clear rights on context sources.
Embedding
Vector representation of text or features for similarity search — foundation for RAG and semantic search.
Vector Store
Storage and index for embeddings with similarity retrieval — often coupled to RAG pipelines.
AI Agent
LLM-driven unit that plans and executes tools/steps — needs guardrails and human-in-the-loop.
Fine-Tuning
Further training a base model on your data — don’t forget rights, lineage, and evaluation.
Model Registry
Versioned store of ML models with metadata, stages, and approvals — analogous to a data product catalog.
MLOps
Operating and delivering ML models: CI/CD, monitoring, retraining, and governance along the model lifecycle.
AI Evaluation
Systematic measurement of quality, safety, and regressions for models/agents — before and after deploy.
Prompt Engineering
Crafting prompts and system instructions — replaces neither grounding nor guardrails and rights checks.
Business Term Link
Link business term ↔ technical field/asset — makes catalog understandable for non-engineers.
Grounding
Bind model answers to retrieved/verified evidence — core of RAG against hallucination.
Chunking
Splitting documents into retrieval units — chunk size steers recall and noise.
Technical Term
Metadata entry for column, topic, or API field — often with type, lineage, and owner in catalog.
Catalog Search
Full-text/semantic search across datasets, terms, and dashboards — quality depends on tags and business links.
Vector Search
Similarity search over embeddings — basis for RAG and discovery in metadata/content.
Context Window
Max tokens a model sees at once — limits prompt, history, and retrieved chunks.
Metadata API
API for catalog metadata (assets, lineage, tags) — basis for automation and IDE integration.
Lineage Export
Export of lineage graph for audit, impact analysis, or external tool — not just UI screenshot.
Tool Calling
LLM selects and calls tools/APIs — needs guardrails, auth, and observability.
Impact Report
Report of which reports/pipelines are affected by schema or term change — from lineage/catalog.
Model Card
Standardized docs on purpose, data, limits, and risks of a model — governance artifact.
Data Drift
Shift in input distributions vs. training — needs monitoring and retraining.
Concept Drift
Change in the input→target relationship — model performance drops despite “same” features.
LoRA
Parameter-efficient fine-tuning — smaller adapters instead of full retrain.
Structured Output
Model answer in a fixed schema (JSON etc.) — eases automation and validation.
DMBOK
Data Management Body of Knowledge (DAMA): shared professional canon for governance, metadata, quality, security, and lifecycle — at Binom bridged via the 8 pillars, not a chapter-by-chapter clone.
Model Context Protocol
Open protocol so assistants connect tools/data sources in a standard way — governance must whitelist allowed tools.
Agentic Workflow
An LLM plans and executes steps with tools — needs guardrails, observability, and clear stop criteria.
Reranker
Second stage after retrieval: re-scores candidates for relevance — costlier than pure vector search, often more precise.
Chain of Thought
Prompt technique where the model emits intermediate steps — can improve reasoning, but leaks thinking traces.
RLHF
Fine-tuning with human feedback/reward — steers behavior, does not replace domain guardrails.
Few-Shot Prompting
Prompt includes a few examples of the desired format — steers output without fine-tuning, consumes context window.
Jailbreak
Attempt to bypass a model’s safety/policy boundaries — guardrails and monitoring must resist it.
Vector Database
Stores embeddings and searches by similarity — foundation for RAG, not a substitute for relational truth.
Temperature
Controls randomness of token sampling — low = more deterministic, high = more creative/error-prone.
AI Act
EU regulation with risk-based duties for AI systems — governance and documentation become mandatory, not optional.
CDMP
DAMA International certification for DMBOK body knowledge — shared professional language; practical start at Binom via pillars, hub, and tools.
Multi-Agent System
Multiple specialized agents coordinate tasks — needs clear roles, shared state, and stop rules.
Eval Harness
Repeatable test battery for LLM/agent quality (golden sets, scorers) — without evals, “better” is only a feeling.
Red Teaming
Deliberately attacking models/agents (jailbreak, injection, bias) — finds gaps before production users do.
Zero-Shot Prompting
Task with no examples in the prompt — cheap on context, often weaker than few-shot for format tasks.
System Prompt
High-priority steering instruction (role, rules) — should be versioned and reviewed like code.
Prompt Template
Reusable prompt structure with placeholders — maintain centrally instead of copy-paste across ten services.
Semantic Caching
Caches answers for semantically similar queries — saves tokens, risks stale/wrong hits without TTL/policy.
Groundedness
Claims are supported by provided sources — a metric against hallucinations in RAG.
Model Routing
Chooses model/route per request (small/large, specialized) by cost, latency, and risk — governance must whitelist routes.
AI Observability
Traces, metrics, and evals for prompts/agents (latency, cost, quality, safety) — without telemetry, no operations.
Agent Memory
Persistent context for agents across sessions — needs retention, isolation, and deletion paths.
DCAM
Maturity/assessment model for data management programs — measures capability; Binom pillars and advisor help find gaps and artifacts, but do not replace a DCAM assessment.
Speculative Decoding
Small model proposes tokens, large model verifies — cuts latency, needs compatible model pairs.
Prompt Chaining
Output of one prompt step feeds the next — needs clear interfaces and per-step evals.
LLM Gateway
Central proxy for LLM calls (auth, routing, logging, limits) — governance point before the model.
Chunking Strategy
How documents are split for RAG (size, overlap, structure) — bad chunking kills retrieval.
Hybrid Search
Combines keyword/BM25 and vector search — better for exact IDs and semantic paraphrases.
Context Compression
Shrinks context before the model (summaries, top-k) — saves tokens, risks information loss.
Embedding Model
Model that maps text to vectors — model choice affects retrieval quality and cost.
Cross-Encoder
Scores query-document pairs jointly — more precise than bi-encoder, costlier at large candidate sets.
Tool Schema
Machine-readable description of allowed tool parameters — without schema, models hallucinate arguments.
Agent Planner
Component that plans steps/tools before execution — without planner, agent becomes ReAct roulette.
Artifact
Reusable, versionable output of governance or pipeline work — e.g. schema, policy draft, incident record, or catalog form; semantics stay stable even when the tool changes.
Agent Executor
Runs planned tool calls and collects results — needs timeouts, retries, and audit log.
Output Moderation
Filters/classifiers on model output (toxicity, PII, policy) — input guard alone is not enough.
Human Feedback Loop
Humans rate/correct model outputs for training or live escalation — needs clear rubrics.
Content Estate
Inventory and operating boundary of all collaboration/content systems (mail, sites, files, chat) before they feed AI or analytics.
Evidence Pack
Repeatable pack of policies, exports, and contextual screenshots with owners — for audit, exit, or incident, not a loose screenshot archive.
Scope Drift
Gradual expansion of purpose, data scope, or use beyond contract/review — common with partners and AI pipelines.
Business Function
Organizational function (sales, finance, HR, IT, engineering, ops, marketing, partner) as a domain entry for pain points, controls, and matching stories — complements roles, does not replace them.
Knowledge Check
Hub for series quizzes, glossary buzzword quiz, and learning progress — checks understanding from stories and terms without a separate LMS.
Story Quiz
Embedded quiz block in story Markdown (quiz fence) — checks a story or series core points and stores progress locally.
RAG Eval
Operational assessment of retrieval vs answer quality (relevance, grounding, hallucination) — distinct from pure model eval.
Sanctioned AI
Officially approved AI use with owner, policy, eval, and HITL — counterpart to shadow AI.
AI Acceptable Use (AUP)
Path-bound rules for which AI use is allowed or forbidden, including escapes and escalation — bound to sanctioned-paths-v1, not a slogan PDF.
AI Impact Pack (FRIA)
Operating artifact for impact/fundamental-rights assessment before high/limited release — a gate, not just a document pile.
AI Red-Team Pack
Recurring adversarial test pack (jailbreak, prompt injection, tool abuse) as a pre-prod gate for sanctioned paths.
Rights Register (AI)
Register that separates license/TDM/copyright flags for retrieval, training, and redistribution — catalog gates instead of implicit assumption.
Embedded AI Surface
AI capability embedded in IDE, BI, CRM, mail, or tickets — needs inventory, scope, and scorecard like a sanctioned path.
GPAI Deployer
Role that puts an AI system into service or uses it — distinct from the GPAI provider; deployer duties are not replaced by provider docs.
Ontology
Agreed meaning model (entities, relations, terms) — needs a sync contract when catalog and platform ontology diverge.
KPI
Measure of goal attainment — needs owner, definition, grain, and time logic, or it becomes a vanity metric.
DSAR
Data-subject request for access or other rights — needs discoverable systems and auditable deadlines.
ROPA
Record of processing activities — purposes, categories, recipients, and retention must stay current and auditable.
DORA
EU regulation on digital operational resilience for finance — ICT risk, testing, and third-party oversight.
NIS2
EU directive on cybersecurity and incident reporting for essential and important entities.
GPAI
General-purpose AI models with broad applicability — provider vs deployer duties differ under the EU AI Act.
Mob Programming
Whole team works on one problem (one driver) — strong for complex pipeline/model topics.
Incident Retrospective
Structured review after incident — focus on system improvement, not blame assignment.
Sprint Review
Demo and feedback at sprint end — for data teams often “show new dataset/metric”, not just tickets.
Capacity Forecast
Forecast of available team capacity vs. backlog — prevents overpromising on data roadmap.
Toil Reduction
Automating repetitive ops (manual reruns, ticket-driven fixes) — SRE principle for data platform teams.
High-Risk AI
AI systems with elevated risk under the AI Act — requiring risk management, data governance, and documentation.
Fairness
Assessing and controlling unequal model impacts across groups — criteria and thresholds must be defined upfront.
Adverse Impact
Demonstrable disadvantage to a group from a decision or model — a gate before production use.
Bias
Systematic skew in data or model behavior — requires measurement, mitigation, and documented residual risk.
Explainability
Traceability of model outputs for decision-makers and affected people — method must match the risk class.
MQL
Marketing-qualified lead by agreed criteria — definition and sales handoff must be contractually clear.
LTV
Expected net value of a customer over the relationship — grain, discounts, and churn assumptions must be documented.
Churn
Loss of customers or revenue in a period — numerator, denominator, and reactivation logic must be unambiguous.
Forecast
Forward-looking estimate of metrics — version, assumptions, and separation from budget must be versioned.
Budget
Approved plan value for a period and org unit — must not be confused with forecast or actuals.
Ledger
Accounting system of record for actuals — management accounting and analytics must respect the boundary.
Headcount
Count of people or FTE — as-of date, employment status, and org assignment must be defined.
Foundation Model
Large pretrained model used as a base for many downstream apps — often GPAI under the AI Act.
Citation
Traceable source attribution for generated claims — a gate for BI copilots and compliance-sensitive answers.
Copilot
Embedded AI assistant in productivity tools — needs allowed paths, context bounds, and output gates.
Retrieval Pipeline
Chain index → retrieve → rerank → context for LLM — quality and latency depend on each step.
Vector Chunk
Text segment plus embedding in vector store — chunk size affects recall and hallucination risk.
Semantic Search Index
Index for similarity search on embeddings — core of catalog search and enterprise RAG.
Query Expansion
Broadens user query (synonyms, LLM) before retrieval — helps with domain-specific terms.
Function Calling API
LLM API layer for structured tool calls — basis for agents with database and API actions.
Index Rebuild
Full rebuild of search/vector index — needed on embedding model change or chunk policy shift.
Threshold Alert
A message when a threshold is crossed — with an owner and a next step. Not a mini-dashboard in the mail and not a substitute for the close dashboard.
Data Export
A file or feed for a named purpose (close, audit, upstream system). Not an interactive dashboard and not a chart collection as consolation.
Corridor (follow-on work)
What domain A leaves so domain B can join: keys, grain, artifact names, and one A per decision type — not a central glossary.
Join Pack
Operating artifact `join-pack.md`: freeze, leave-open, brief-customer, neighbor countersign, and join test. Practice after the first review — not Sales.
Freeze / leave-open
What must hold for every follow-on (keys, grain, filenames, one A) versus what stays local or deliberately undecided (semantic layer, harvest, “one number”).
Project moment
The week in the customer project (CRM go-live, close, mart cut) where the governance SKU attaches — not the Functions hub card.
Find, do not expose
Three layers of internal visibility: work need-to-share, employment records need-to-know, directory with a mandatory core and a choice. One transparency policy for everything creates silos or glass employees.
Visibility contract
Operating artifact `visibility-contract.md`: three layers, work allowlist, record gate, mandatory vs optional core, preference record, revocation through the index, kill line. Practice — HR/privacy countersign.
Directory preference
Steering of optional directory fields (photo, bio, skills, audience). Not consent and not a legal basis for people analytics or record access. Revocation must hit UI and index.
Employment record path
The only allowed channel for contracts, pay, health, performance, and IDs: HRIS plus a named record site. Search, Copilot, and general marts do not inherit the record. Shadow copies count.
Serial grain
One row per physical device with a stable ID — not quote SKU quantity and not a CMDB dump.
Install-base
Shipped, still-active devices with location and contract — not the CRM product catalog and not quote quantity.
Deal registration
Protection claim in the vendor portal with a source ID. Registered is not a CRM pipeline stage and not a booking.
Rebate vs booking
A manufacturer or distributor payment (rebate/SPIFF) is not a booking, not hardware margin, and not maintenance ARR.
RMA
Return or replacement on the serial — not on the ticket alone and not on account 360.
Field-service SLA
Service clock on the serial (and site), not on the account. A ticket without a serial is not an SLA case.
Medallion
Layer model for raw, cleaned, and consumable data — a storage pattern, not a full governance architecture and not a substitute for owner, contract, and purpose.
Codetermination
DACH C-gate: the works council is Consulted on employee data, not the Data Owner. Evidence before processing; a return path and scope are mandatory.
Trust claim
Which consumer decision becomes more reliable in 90 days, and how the person notices it. Catalog coverage does not count. No trust claim, no Placement.
Placement
Default Field-Sales SKU: partial scope on a live initiative without owning the stack. A pilot needs a consumer; a platform lane needs a governance owner.
SKU
One deliverable service cell: Placement, Pilot, platform lane, or handover — one, not four. Sales does not design Ranger, dbt, or a semantic layer as a giveaway.