Zum Inhalt springen
Search the hub

Glossary

Shared vocabulary for data governance — definitions with links into stories, tools, and learning paths.

Funny Meeting Bingo
833
Data & products Authoritative Source The decided source of truth for a clear purpose — with owner and evidence. Related to SSOT; stresses purpose-binding over “one system for everything”. Roles Control Owner Accountable for the effectiveness of a concrete control (e.g. access review, masking, deletion gate) — often Security/Privacy/Platform, not the Data Owner of business meaning. AI & ML FRIA Fundamental-rights impact assessment for high-risk AI: purpose, affected people, risks, mitigations, and evidence — before go-live and on material change. Data & products Identity Resolution Matching and separating party/customer identities across sources (match, merge, survivorship) — with purpose limits and owners. Process Minimum Viable Governance Smallest workable governance setup for SMB/mid-market: clear decision rights, one cohort product, measurable cadence — without role sprawl and without a tool landscape as a substitute for accountability. Process Operating Cadence Recurring rhythm for reviews, approvals, intake, and adoption measurement. Cadence makes governance operational — instead of one-off project slides. Privacy & protection Re-identification Risk that anonymized or synthetic data can reveal individuals again. Needs gates, tests, and release approval. Privacy & protection Subprocessor Further processor in the processing chain. Governance needs inventory, contract chain, change notification, and exit/evidence paths. Data & products Survivorship Rules for which source attribute wins conflicts in a golden record — with owner, evidence, and review. Without survivorship, identity resolution stays unresolved conflict. Privacy & protection TIA Assessment of third-country transfers: legal basis, recipients, risks, supplementary measures, and review cadence — complements SCCs, does not replace them. Roles Data Steward Operational ownership for definition, quality, and use of a data domain — not documentation alone. Roles Data Owner Business decision-maker for purpose, access rules, and approvals of a data product or domain. Roles Data Architect Owns grain, model consistency, and contracts — so domains and marts fit together without replacing owner or steward. Roles Data Custodian Technical custody of systems and storage — access upkeep, backups, runtime — usually platform/IT, not business definition. Roles Data Consumer Uses data products for decisions or reports — raises quality issues but does not alone decide definition and access. Roles Data Product Owner Owns product lifecycle, priorities, and consumer value — distinct from the domain Owner and the Steward. Roles Analytics Engineer Builds trusted transforms, tests, and docs (often with dbt] between platform and BI. Roles Governance Center of Excellence Central enablement for standards, cadence, tooling, and cross-domain escalation. Roles Governance Lead Accountable for the governance operating model and sponsor evidence — not every ticket. Roles Data Engineer Builds and runs pipelines, storage layers, and integration patterns — delivering reliable inputs for stewardship and analytics. Roles BI Developer Implements semantics and visualization in BI tools — ideally on certified datasets, not by inventing business logic inside reports. Roles Technical Owner Owns technical operability of a data product or pipeline — incidents, deployments, tooling — alongside the business Data Owner. Roles Executive Sponsor Provides mandate, priority, and escalation power for governance or platform initiatives — without day-to-day stewardship work. Roles CDO Leadership role for data strategy, mandate, and enterprise prioritization — not day-to-day stewardship itself. Roles Platform Ops Operates orchestration, environments, and platform SLAs — keeps the pipeline machinery running. Roles Citizen Developer Business user who builds reports/automations themselves — needs governed semantics, or shadow IT follows. Roles Stakeholder Person or group with a stake in data decisions — needs clear R/A/C/I, or role sprawl follows. Roles Mandate Formal authorization for a role or CoE to drive decisions and enforce policies. Roles Chief Data Officer Executive for data strategy, governance, and value — sets priorities and decision rights, does not write every KPI personally. Roles Analytics Translator Translates business questions into analytical problems and results back into decisions — bridge between business and analytics. Roles ML Engineer Puts models into production: serving, monitoring, retraining pipelines — not notebook experiments alone. Roles Platform Engineer Builds self-service paths (CI, environments, observability) for teams — cuts toil, does not own every dashboard. Roles Data Scientist Finds patterns and models from data for decisions — experiments, but does not replace platform pipeline ownership. Roles Prompt Engineer Designs prompts, evaluations, and guardrails for LLM use — focuses on behavior, not infrastructure. Roles AI Product Manager Prioritizes AI features by value, risk, and measurability — ties business outcomes to model/eval metrics. Roles Head of Data Leads the data team operationally (priorities, delivery, people) — often under a CDO, closer to day-to-day than pure strategy. Roles AI Engineer Builds production AI features (RAG, agents, evals, serving) — bridge between research notebooks and platform. Roles Domain Owner Owns a business domain including data products and decisions — not ticket prioritization alone. Roles Data Governance Manager Coordinates governance program, policies, and rollout — sets the frame, does not own every dataset personally. Roles Analytics Lead Leads analytics team and priorities (backlog, quality, adoption) — bridge between BI/DS and business. Roles DataOps Engineer Automates and runs data pipelines (CI/CD, monitoring, recovery) — ops for data, not infra alone. Roles Compliance Analyst Translates regulation into auditable requirements for data/BI — not checklists without technical context. Roles FinOps Analyst Analyzes and steers cloud/data costs (showback, rightsizing) — saves money, does not replace architecture. Roles Information Architect Structures information and taxonomies for findability and governance — bridge between content and data catalog. Roles BI Analyst Builds reports/dashboards and translates questions into models — closer to business than pure pipeline engineering. Roles Reporting Analyst Specialist for recurring, governed reports — focus on consistency and distribution, not ad-hoc exploration. Roles Chief Analytics Officer Executive for analytics strategy and ROI — owns insight adoption, not just tool budgets. Roles Data Engineering Manager Leads data engineering teams and prioritizes platform vs. feature work — people + delivery, less hands-on SQL. Roles Governance Analyst Operationalizes policies, workflows, and metadata standards — bridge between legal/compliance and engineering. Roles Privacy Engineer Implements privacy-by-design in pipelines and APIs — technical execution of consent, retention, and access. Roles Security Engineer Builds secure defaults into data platforms (IAM, encryption, secrets) — not just closing pen-test tickets. Data & products Data Product Consumable, versioned data offering with a clear audience, contracts (SLA/SLO), ownership, and documented quality. Data & products Grain The smallest business statement unit of a table or mart (e.g. one order, one day, one contract]. Data & products Semantic Layer Executable meaning layer for reusable metrics, dimensions, filter logic, and aggregation rules above physical tables. It defines how a claim is calculated; a dashboard presents the claim, and a catalog makes it discoverable. Data & products Semantic Model Tool-bound semantic artifact (e.g. Power BI model, Qlik app model]. Data & products Data Domain Bounded business area of ownership, products, and decision rights. Data & products Metric / KPI Store Technical or logical store for governed metric definitions, versions, owners, tests, and approval status. A metric store can be part of a semantic layer, but it is not the visual presentation in a dashboard. Data & products Data Product Certification Explicit trust status that a product meets contract, DQ, and ownership bars. Data & products Data Product Versioning Explicit versions and compatibility rules for product interfaces. Data & products Breaking Change Interface, schema, or semantics change that invalidates existing consumers. Data & products SLA / SLO Service promises (SLA] and measurable reliability targets (SLO], e.g. freshness. Data & products Reverse ETL / Activation Push curated warehouse data back into operational tools. Data & products Data CI/CD Automated test and deploy of models, contracts, and quality checks. Data & products Trusted Metrics Metrics with a clear contract, owner, grain, and versioning — so reports consume the same truth. Architecture Medallion Architecture Bronze/Silver/Gold as technical zones — useful labels, not a full logical warehouse model. Data & products KPI Contract Agreement on definition, grain, filters, owner, and breaking-change rules for a metric — before the dashboard. Data & products Single Source of Truth Authoritative source / SSOT: the decided, evidenced source of truth for a clear purpose (e.g. a metric or a master record). SSOT is a governance decision with an owner and evidence — not a slogan and not a mandate to store everything in one physical system. Data & products Master Data Shared core entities (customer, product, location) with ownership and matching — foundation for golden records. Data & products Reference Data Controlled value lists and codes (country, currency, status) — small, stable, often underestimated in governance. Data & products Data Literacy Ability to read, question, and use data meaningfully — prerequisite for self-service without chaos. Architecture Landing / RAW Layer Ingest as received: preserve source payload and load identity with minimal semantic change. Data & products Deprecation Planned retirement of a data product or measure — including deadline, replacement, and consumer communication. Data & products Data Product Lifecycle Phases from intake through build, operate, versioning, to deprecation of a data product. Data & products Subject Area Business topic block (e.g. finance, customer) — often the scope for domains, marts, and stewardship. Data & products Producer / Consumer Relationship between who delivers a data product and who consumes it — contracts and SLAs clarify both sides. Data & products MVP Smallest usable slice of a data product — vertical slice instead of big bang, with a measurable outcome. Architecture Conform / Standardized Layer Technical standardization and validation before business identity integration. Data & products Data Mart Business-scoped analytical dataset — often the business data product layer. Data & products Dark Data Stored but unused/undocumented data — cost, risk, and ROT candidate. Data & products ROT Redundant, obsolete, or trivial data — cleanup saves cost and privacy surface. Data & products Data Strategy Target picture and priorities for data products, capabilities, and governance — not the tool list. Data & products Data Roadmap Timed sequence of outcomes and data products — tied to value streams, not tools. Data & products Domain Ownership Ownership follows the domain, not only a central team — core of mesh and federated governance. Data & products First-Party Data Customer data collected directly by the company — often most valuable, but heavily regulated. Data & products Third-Party Data Data from external providers — scrutinize contracts, provenance, and purpose limitation. Data & products Zero-Party Data Preferences/intent consciously shared by the user — keep consent and purpose clear. Architecture Integrated Core Shared business entities, relationships, and history across sources. Data & products Crown Jewel Data Highest-value/highest-sensitivity assets — maximum controls, monitoring, and ownership. Data & products Data as a Product Datasets are run like products: clear owner, SLA, versioning, and user promise — not an ops by-product. Architecture Mart / Business Data Product Layer Purpose-bound facts, dimensions, and KPI bases for a defined consumer purpose. Data & products Data SLA Measurable promise for a data product (freshness, availability, correctness) — without an SLA, “important” stays opinion. Data & products Dataset Versioning Traceable versions of a dataset (schema + content) — enables reproducibility and rollback. Data & products Trusted Data Source Authoritative source for a business entity — downstream may copy, not contradict without a contract. Data & products Master Data Management Discipline to provide consistent, versioned master data (customer, product, place) — not spreadsheet matching alone. Data & products Semantic Contract Binding agreement on meaning and calculation of metrics/entities — not column types alone. Data & products Operational Data Store Near-ops, often integrated copy for operational queries — not a historical warehouse and not a lakehouse substitute. Data & products Metric Layer Central definition of metrics (formula, grain, owner) — prevents per-tool “revenue” variants. Data & products Data Product Canvas One-page sketch of users, value, interface, quality, and owner for a data product — before building pipelines. Data & products Domain Data Product Data product of a business domain with a clear contract — cross-domain use only via published interfaces. Data & products Data Sharing Agreement Governs purpose, rights, quality, and liability when sharing data between parties — not just a SharePoint link. Data & products Platform as a Product Internal platform is run like a product (roadmap, UX, SLAs for teams) — not as ticket helpdesk. Architecture Consumption Contract Tool-specific access and interpretation of a governed product (views, semantic models, extracts]. Data & products Self-Serve Platform Teams can provision data products/pipelines without gatekeeper tickets — guardrails instead of approval queues. Data & products Data Mesh Principle Guardrails for decentralized data products (domain ownership, self-serve, federated governance) — not a tool product name. Data & products Data Contract Testing Automated tests against the published data contract — if producer breaks, build fails, not the dashboard. Data & products Data Portability Data subjects can receive/transfer data in machine-readable form — technical export paths must exist. Data & products Data Product Specification Binding description of interface, SLA, schema, and owner — where the canvas becomes operational. Data & products Data Democratization Broader, safe data access for non-engineers — needs guardrails, not “everyone may access everything”. Data & products Data Entitlement Right to specific datasets/columns for role/person — enforceable technically, not policy PDF alone. Data & products Data Product Roadmap Timeline for data product features and dependencies — links business outcomes to pipeline/model work. Data & products Data Monetization Turns data deliberately into revenue/efficiency — needs legal basis, quality, and measurable value. Data & products Data Usage Metrics Metrics on queries, downloads, and dashboard use — shows value and hotspots, not just storage size. Data & products Data Consumption Pattern Typical read pattern (batch vs. streaming, OLAP vs. API) — drives caching, partitioning, and cost. Data & products Data Product Discovery Users find fitting data products in catalog/marketplace — without discovery, mesh stays architecture folklore. Data & products Data Provenance Tracking origin and transformations of a dataset — foundation for trust and audit. Data & products Federated Computational Governance Global standards enforced as automated checks/policies — governance as code, not CoE slides alone. Data & products Data Provenance Chain Continuous chain of sources, jobs, and releases to the report — stronger than isolated lineage snapshots. Architecture Lakehouse Open table formats plus warehouse-style governance on shared lake storage. Architecture Data Mesh Domain-owned data products with federated governance, not a tooling checklist. Architecture Simplest Viable Architecture Least unnecessary complexity that still meets the requirement — not fewest boxes. Architecture Modern Data Warehouse Governed path from source to products and semantics — not just a cloud vendor rename. Architecture Greenfield Build-from-scratch warehouse without strangling a legacy estate. Architecture Brownfield Modernize an existing warehouse or estate in place. Architecture Strangler Pattern Incrementally replace legacy paths by routing new traffic to the modern stack. Architecture Vertical Slice End-to-end thin path (source → product → consumer] before horizontal platform sprawl. Architecture Hybrid Cloud On-prem and cloud coexistence as a deliberate architecture, not a temporary defect. Architecture ELT Load first, transform in the platform (typical dbt pattern]. Architecture Orchestration Scheduling and dependency management across pipelines and quality gates. Architecture dbt SQL-first transform, test, and documentation framework in the warehouse/lakehouse. Architecture Microsoft Fabric Unified analytics SaaS (OneLake, Lakehouse/Warehouse, Power BI, Purview touchpoints]. Architecture OneLake Fabric’s single logical lake storage plane for Lakehouse and Warehouse assets. Architecture Apache Spark Distributed compute for batch/streaming on lakes and warehouses — often behind notebooks and lakehouse engines. Architecture Notebook Interactive compute surface for exploration and jobs — production-ready only with versioning, tests, and ownership. Architecture SQL Warehouse SQL compute endpoint on lake/warehouse data — separates storage from elastic query compute. Architecture Warehouse-Native Transformation Transformations run inside the warehouse/lakehouse (ELT), not on a separate ETL server. Architecture Parallel Run Old and new pipeline running together — compare before cutover to catch regressions and trust gaps. Architecture FinOps Cost accountability for cloud/compute use — tags, budgets, and idle cleanup belong in platform governance. Architecture dbt Seed Versioned CSV/flat data in dbt — typical for reference data and small lookups. Architecture dbt Snapshot dbt mechanism for historized snapshots of source tables — related to SCD2/as-was history. Architecture dbt Exposure Declared downstream use (dashboard, ML, app) in dbt — makes impact and ownership visible. Architecture Delta Lake ACID table format commonly underlying lakehouse cores and DQ result stores. Architecture dbt Macro Reusable Jinja logic in dbt — e.g. for RAW generation, tests, and naming conventions. Architecture Materialization How dbt builds a model (view, table, incremental, ephemeral) — drives cost, latency, and downstream contracts. Architecture Incremental Model dbt materialization that processes only new/changed rows — needs clear unique keys and a late-arrival strategy. Architecture Ephemeral Model dbt model without a persisted object — inlined as a CTE into downstream models. Architecture ETL Transform before load — classic pattern; often replaced or hybridized by ELT in the warehouse today. Architecture Transformation as Code Transformations versioned in the repo with review and CI — instead of click-ETL and undocumented jobs. Architecture Cutover Moment consumers switch from old to new stack — needs parallel run, rollback, and clear contracts. Architecture Parquet Columnar file format for analytical workloads — standard in lakes and many warehouse exports. Architecture Apache Iceberg Open table format for lakehouses — ACID, schema evolution, and time travel on object storage. Architecture Apache Ranger Policy engine for fine-grained access control on Hadoop/lakehouse workloads — often paired with Atlas for audit and lineage. Architecture QVD Qlik’s optimized columnar extract format for reusable load layers. Architecture Hive Metastore Central metadata service for tables, partitions and schemas in Hadoop/lakehouse stacks — foundation for Spark, Hive and Ranger policies. Architecture ODS Integrated operational staging close to source systems — often a precursor to the warehouse, not a mart replacement. Architecture Apache Atlas Open-source governance and lineage catalog for the Hadoop ecosystem — classification, tags and audit events for Ranger-protected assets. Architecture Data Lake Storage for raw and semi-structured data at scale — without warehouse semantics; often a precursor or part of a lakehouse. Architecture Apache NiFi Visual data movement and integration framework — flows need ownership, provenance and access control like any pipeline. Architecture Data Vault Modeling pattern with hubs, links, and satellites for auditable historization — alternative/complement to a dimensional core. Architecture AWS Lake Formation AWS service for lake governance — central permissions, LF tags and data location management across S3, Glue and analytics engines. Architecture System of Record The authoritative operational system for an entity or attributes — not the same as an analytical single source of truth. Architecture AWS Glue Serverless ETL and metadata catalog in AWS — crawlers, jobs and catalog tables as the base for Lake Formation and Redshift/Spectrum. Architecture Data Fabric Architecture approach with metadata, virtualization, and integration across distributed sources — often a marketing umbrella beside mesh/lakehouse. Architecture Apache Kafka Distributed event streaming — topics, schemas and ACLs need governance analogous to batch pipelines (registry, ownership, PII tags). Architecture Logical Architecture Layers, contracts, and responsibilities independent of specific tooling — the foundation before physical stack choice. Architecture Kafka ACL Access control on Kafka resources (topics, consumer groups, cluster) — keep authority and enforcement separate from schema and contract governance. Architecture Streaming Continuous event/record processing — only where latency and cost justify the complexity. Architecture Confluent Commercial Kafka ecosystem (Platform/Cloud) with Schema Registry, Connect and governance features — same concerns as open-source Kafka, often with a stronger managed control plane. Architecture Near Real-Time Latency of seconds to a few minutes — costlier and more complex than batch; only where decisions truly need it. Architecture Batch Processing Scheduled, periodic processing of large volumes — the default for most warehouse and mart loads. Architecture Apache Hadoop Distributed storage and compute ecosystem (HDFS, YARN and related services) — foundation for many on-prem/hybrid lakehouse landscapes including Cloudera CDP. Modeling Dimensional Modeling Facts and dimensions organized for analytics grain and reuse (Kimball-style]. Architecture HDFS Distributed file system in the Hadoop stack — blocks, replication and paths are typical Ranger/audit resources. Architecture Jinja Templating language behind dbt macros and models — powerful, but unreadable without standards. Architecture Backfill Loading historical periods after the fact — needs idempotency, watermarks, and clear cutover rules. Architecture YARN Cluster resource manager for Hadoop workloads — queues and compute limits belong to the operating boundary alongside data access. Architecture Apache Hive SQL engine and table model on the lake — policies and ownership often attach to Hive tables while the Hive Metastore holds technical truth. Architecture Watermark Progress marker for incremental loads (timestamp/id) — set wrong, you get gaps or duplicates. Architecture Apache Ozone Object storage for Hadoop/CDP landscapes — scalable alternative or complement to HDFS with its own resource boundaries for policies. Architecture Idempotent Load Re-running the same job does not change the result uncontrollably — mandatory for retries and backfills. Architecture Full Load Full reload of a table/partition — simple, expensive; often only for bootstrap or repair. Architecture Kerberos Authentication protocol for cluster identities — in CDP/Hadoop often the bridge from corporate directory to service principals and Ranger. Architecture Apache Knox Perimeter gateway for Hadoop/CDP services — centralizes authn/authz at the cluster edge and decouples clients from internal endpoints. Architecture Soft Delete Mark a record deleted instead of physically removing it — helps audit and CDC, complicates queries. Architecture Cloudera CDP Cloudera Data Platform for hybrid/multi-cloud and on-prem analytics — SDX bundles Shared Data Experience (Ranger, Atlas, HMS) as a governance surface, but does not replace a business operating model. Architecture Time Travel Querying a prior table version (lakehouse/warehouse) — useful for audit and rollback, costs retention. Architecture Cloudera SDX Shared Data Experience in CDP — shared security, governance and metadata services (including Ranger, Atlas, Metastore) across workloads. Architecture Zero-Copy Clone Clone of a table/DB without duplicating data — ideal for dev/test when lifecycle and rights are clear. Architecture Apache Impala Massively parallel SQL engine on Hive/HDFS/Ozone data — often the interactive analytics path in CDP alongside Spark and Hive. Architecture Partition Pruning Query reads only relevant partitions — requires clean partition keys and filters in contracts. Architecture Apache ZooKeeper Coordination service for distributed systems — quorums, leader election and config for Kafka, HDFS NameNode HA and other CDP services. Modeling Star Schema Fact table surrounded by denormalized dimensions. Architecture Cloudera Manager Operations console for Cloudera clusters — service lifecycle, config and health; a custodian tool, not business authority. Architecture Predicate Pushdown Filters execute as close to storage/engine as possible — less I/O, faster marts and BI. Architecture Apache HBase Distributed wide-column store on HDFS — its own namespace/table resources for Ranger and often tied to operational workloads. Architecture Compaction Merging small lake files and cleanup — keeps scan cost and time travel in check. Architecture Apache Hudi Lakehouse table format with upserts/incremental processing — alternative/complement to Delta and Iceberg. Architecture LDAP / Corporate Directory Enterprise directory for identities and groups — mapping to Ranger policies and service accounts is the critical governance link, not the group name alone. Architecture Avro Row-oriented serialization format with schema — common in streaming and schema registries. Architecture Schema Evolution Planned schema change without uncontrolled breaks — needs compatibility rules and contracts. Architecture Workflow Orchestrator System for scheduling, dependencies, and retries of data jobs — not the same as transformation (dbt). Architecture Ingestion Tool Connector/ELT tool loading from SaaS/DBs into landing — replaces neither modeling nor contracts. Architecture Protocol Buffers Compact binary schema format — common in APIs and event streams with strict evolution. Architecture ORC Columnar format for analytical workloads — Parquet alternative, strong in Hive ecosystems. Modeling Fact Table Measurable events or transactions at a declared grain. Architecture Apache Arrow In-memory columnar format for fast exchange between engines — bridge across Spark, DuckDB, BI. Architecture DuckDB Embedded analytical DB for local/edge SQL workloads — great for exploration, not automatically an enterprise hub. Architecture Log-Based CDC Change data capture from database logs — low source impact, needs privileges and lag monitoring. Architecture Exactly-Once Processing promise with no duplicates and no loss — expensive; often effective exactly-once via idempotency is enough. Architecture At-Least-Once Every event arrives at least once — duplicates possible; downstream must be idempotent. Architecture Replay Reprocessing historical events/loads — needs idempotency and clear watermarks. Architecture Data Share Governed exchange of datasets across boundaries — without uncontrolled exports. Architecture Clean Room Controlled environment for joint analysis without raw data exchange — privacy by design. Architecture Blue/Green Deployment Two parallel environments; switch only after validation — related to parallel run/cutover. Modeling Dimension Table Descriptive context (customer, product, calendar] joined to facts. Architecture Canary Release New release first for a small traffic/user share — limit risk before full cutover. Architecture Feature Flag Runtime switch for features/pipelines — decouples deploy from release, needs governance. Architecture Outbox Pattern Publish events reliably from the same transaction as the state change — against message loss. Architecture Infrastructure as Code Infra versioned and reviewable in the repo — foundation for reproducible data platforms. Architecture GitOps Git as source of truth for deploy state — pull-based, auditable, good for platform standards. Architecture CQRS Separation of write and read models — useful under different write/read loads, increases complexity. Architecture Saga Sequence of local transactions with compensation instead of distributed 2PC — for multi-service flows. Architecture Polyglot Persistence Multiple storage technologies per use case — needs clear ownership and integration contracts. Architecture Change Data Feed Table change stream from lakehouse formats — downstream reads inserts/updates/deletes incrementally. Modeling Conformed Dimension Shared dimension meaning and keys reusable across marts and domains. Architecture Z-Order File clustering by multiple columns — improves skip/pruning for mixed filters. Architecture Liquid Clustering Flexible clustering strategy in modern lakehouses — less rigid partitioning. Architecture Vacuum Cleanup of unreferenced files after time-travel retention — cost and compliance lever. Architecture Shallow Clone Clone that shares metadata and does not copy data — fast, but lifecycle-coupled. Architecture Deep Clone Clone with an independent data copy — costlier, but decoupled from source lifecycle. Architecture Data Marketplace Catalog/marketplace to find and consume data products — needs certification and contracts. Modeling SCD (Slowly Changing Dimension] Pattern family for how dimension attributes change over time. Architecture Zero-ETL Architecture promise: analytics close to the source without classic batch ETL copies — governance and latency still matter. Architecture Data Virtualization Queries across distributed sources without physical consolidation — fast to integrate, often costly for performance and lineage. Architecture Lambda Architecture Parallel batch and speed layers for correctness plus low latency — duplicated logic is the classic cost. Architecture Kappa Architecture One stream pipeline for realtime and replay instead of separate batch/speed layers — simplifies code, needs a robust event log. Architecture Event Sourcing State is derived from an immutable event sequence — strong for audit, but replay and schema evolution are demanding. Modeling SCD Type 2 Historize attribute changes with effective dating or version rows. Architecture Message Broker Middleware for async messages (topics/queues) — decouples producers and consumers in event architectures. Architecture Event Choreography Services react to events without central orchestrator — flexible but harder to debug than orchestration. Architecture Event Orchestration Central coordinator drives steps and retries across services — easier to monitor than pure choreography. Architecture Saga Orchestrator Component running distributed transactions as sagas with compensating actions — common in microservice stacks. Architecture Warm Standby Standby system is prepared and quick to start — balance between cost and recovery time vs. cold standby. Architecture Active-Active Multiple sites/clusters take traffic in parallel — highest availability, complex for data consistency. Architecture Geo-Replication Replicates data across regions for DR and latency — cost and consistency model must align. Architecture Multi-Region Architecture across multiple cloud regions — affects data residency, failover, and compliance together. Modeling CDC (Change Data Capture] Capture source inserts, updates, and deletes for incremental loads. Modeling Delta / Incremental Load Process only changes since last successful run (often via MERGE]. Modeling Surrogate Key System-generated durable key independent of source natural keys. Modeling Natural Key Business or source identifier used for matching and lineage back to origin. Modeling Golden Record Resolved, governed master representation of an entity (e.g. Golden Customer]. Modeling As-Was History Point-in-time view of relationships and attributes as they were at a past date. Architecture Outbox Pattern Writes domain event and DB change in one transaction; a relayer publishes async — avoids lost updates between DB and bus. Architecture Saga Pattern Orchestrates distributed steps with compensations instead of one global DB transaction — failure is part of the design. Architecture Event-Driven Architecture Systems react to events instead of sync call chains — decouples producers/consumers, needs clear event contracts. Architecture Micro-Batch Processes events in short intervals (seconds/minutes) instead of one-by-one — trade-off between latency and throughput. Architecture Idempotency Applying the same request/event multiple times yields the same end state — required with retries and exactly-/at-least-once. Architecture Write-Audit-Publish Write and audit data first, only then publish to consumers — prevents “broken goes live”. Architecture CQRS Read Model Read-optimized model separate from the write model — often denormalized and updated via events. Architecture Data Replication Copies data between systems for latency, resilience, or isolation — needs conflict and lag strategy. Architecture Fan-Out One event/request is delivered to many consumers — scales read load, increases observability need. Architecture Append-Only New entries are only appended, never overwritten — audit and replay get easier; corrections are new events. Architecture Event-Carried State Transfer Event carries enough state for consumers to work without querying the producer — trade-off: larger payloads. Modeling Schema Drift Unexpected source or schema change that breaks contracts or loads. Architecture Strangler Fig Pattern Replace legacy incrementally, route by route — avoid big-bang migration. Architecture Bulkhead Pattern Isolate resources so one failure does not take everything down — e.g. separate pools per tenant/domain. Modeling Referential Integrity Guarantee that foreign keys point to existing parents — central for facts, dimensions, and DQ gates. Architecture Sidecar Pattern Helper process runs beside the main service (logging, mesh, auth) — decouples cross-cutting from app code. Modeling Bridge Table Helper table for many-to-many between facts and dimensions (or among dimensions) without breaking grain. Architecture Service Mesh Infrastructure layer for service-to-service (mTLS, routing, observability) — policy without app changes. Architecture Materialized View Stored query results for fast reads — refresh strategy (incremental/full) is part of the design. Architecture Read Replica Readable copy of the primary DB — offloads analytics; replication lag must be communicated. Modeling Snowflake Schema Dimensional variant with normalized dimension tables — saves redundancy, often costs join complexity vs. star. Architecture Cold Storage Cheap, slow storage tier for rarely used data — retrieval time belongs in the SLA. Architecture Anti-Corruption Layer Translates foreign/legacy model into your domain model — stops legacy language infecting the design. Modeling Junk Dimension Bundles low-cardinality flags/codes into one dimension — keeps the fact table lean. Modeling Degenerate Dimension Dimensional attribute that lives only on the fact table (e.g. invoice number) — without its own dimension table. Architecture Hexagonal Architecture Domain core in the center, I/O via ports/adapters — testable and swappable without bending the core. BI & Reporting DAX Formula and query language for Power BI / Analysis Services — measures and calculated columns in the semantic model. Architecture Ports and Adapters Port defines interface, adapter implements it (DB, API, UI) — core knows no SQL details. Architecture Event Backbone Central event infrastructure domains plug into — scaling and schema governance are shared duties. Modeling SCD Type 1 Dimension change overwrites the current value — no history; simple and often wrong for audit. Architecture Log Shipping Ships transaction logs to standby/replica — classic for DR; latency and failover must be planned. Modeling SCD Type 3 Stores limited prior versions in extra columns — rare; usually SCD2 or as-was is better. Architecture Hot Path Path for latency-critical processing (real-time/near-RT) — costlier, must stay lean. Modeling Role-Playing Dimension Same dimension used multiple times in different roles (order date, ship date) — without duplicating physical tables. Architecture Cold Path Path for batch/archive processing with higher latency — cheaper, for history and heavy aggregation. BI & Reporting Set Analysis Qlik syntax to compute aggregations independently of — or deliberately against — the current selection state. Architecture Backpressure Downstream signals “slow down” so upstream does not overflow — without backpressure, streaming crashes. Modeling Factless Fact Table Fact table without numeric measures — captures events or coverage (e.g. attendance, promotion coverage). Modeling Late-Arriving Data Facts or dimensions arrive after the expected load window — incremental/SCD strategies must tolerate it. Modeling Matching Decision of which source records are the same entity — prerequisite for golden record and MDM. BI & Reporting LookML Looker’s modeling language for views, explores, and measures — business logic as code instead of inside every report. Modeling Historized Entity Entity with a validity timeline (valid from/to) — core of the integrated core and as-was reporting. Modeling Bus Matrix Matrix of business processes × conformed dimensions — planning tool for marts and shared dimensions. Modeling Additive Measure Measure that can be summed across all dimensions (e.g. revenue) — foundation of clean aggregations. BI & Reporting LOD Expression Tableau expressions that control aggregation level independently of the current viz grain. Modeling Semi-Additive Measure Summable across some dimensions, often not across time (balances) — needs snapshot logic. Modeling Non-Additive Measure Not meaningfully summable (ratios, distinct counts) — aggregation must be recalculated. Modeling Transaction Fact Fact per business event (order line) — fine grain, additive measures. BI & Reporting Filter Context The set of active filters/selections under which a measure formula evaluates — the source of many “wrong” KPI numbers. Modeling Periodic Snapshot Fact Fact with a regular status picture (day/month) — typical for balances and semi-additive measures. Modeling Accumulating Snapshot Fact Fact that progresses a process across milestones (apply → approve → close) — rows are updated. Modeling Kimball Dimensional approach with bus matrix and conformed dimensions — bottom-up marts around business process. BI & Reporting Master Measure Reusable, versioned measure definition in the semantic layer — not reinvented per report. Modeling Inmon Top-down enterprise DW with normalized core and dependent marts — contrast to the Kimball bus. Modeling Canonical Model Shared integration model across systems — reduces point-to-point, needs strong ownership. Modeling Many-to-Many Relationship where both sides have multiple partners — in dimensional models often via bridge tables. BI & Reporting Calculation Group Tabular/Power BI pattern to model time intelligence or presentation variants centrally instead of measure explosion. Modeling Cardinality How many partners a relationship has (1:1, 1:n, n:m) — drives joins, grain, and BI filter paths. Modeling Sparsity Many combinations without events — affects storage, aggregations, and BI performance. Modeling Outrigger Dimension Dimension hanging off another dimension instead of the fact — use sparingly. BI & Reporting Self-Service BI Business users build reports themselves — works only with governed semantics and clear consumption contracts, otherwise shadow IT. Modeling Mini-Dimension Small dimension for rapidly changing attributes — relieves large SCD2 dimensions. Modeling Late-Arriving Dimension Fact arrives before the dimension — needs inferred/unknown members and later matching. Modeling Multi-Valued Dimension One fact has multiple dimension values at once — typically solved via bridge/weight factors. BI & Reporting Embedded Analytics Analytics and visuals embedded in operational apps — same contracts and entitlements as standalone BI. Modeling SCD Type 6 Hybrid of types 1/2/3 — current and historical views in parallel, modeling-heavy. Modeling SCD Type 0 Attribute stays unchanged (retain original) — rare; document clearly. Modeling Heterogeneous Dimension One dimension table for different entity types with shared keys — use carefully. BI & Reporting Metric Lineage Traceability from measure/KPI back to sources, transformations, and owners — mandatory for trusted metrics. BI & Reporting Certified Dataset Approved, reviewed semantic dataset for self-service — a badge does not replace owner and contract. Modeling Slowly Changing Dimension Pattern for dimension changes over time (Type 1 overwrites, Type 2 historizes) — without a strategy every trend analysis lies. BI & Reporting DirectQuery BI query mode against the source instead of import — choose freshness vs. performance and governance trade-offs deliberately. BI & Reporting Thin Consumer Interface Reports hold little business logic of their own — they consume measures and contracts from the semantic layer. BI & Reporting Headless BI Semantics and metrics as an API/service — visualization is swappable, definitions stay central. BI & Reporting OLAP Analytical multidimensionality (slice/dice, hierarchies) — today often as tabular/columnar engines instead of classic cubes. BI & Reporting Tabular Model Columnar semantic model (e.g. Power BI / AAS) with relationships, measures, and often import/DirectQuery. BI & Reporting Shadow IT Ungoverned reports, exports, and pipelines outside the platform — symptom of missing semantics and delivery. BI & Reporting Time Intelligence Period comparisons and YTD/MTD logic — central in the semantic layer, not copy-pasted into every report. Quality Data Quality Measurable fitness of data for a purpose — rules, ownership, monitoring, and remediation instead of one-off checks. BI & Reporting Composite Model Semantic model with mixed storage modes — flexibility with higher governance and performance complexity. BI & Reporting Master Dimension Reusable dimension/hierarchy in the semantic layer — analogous to a master measure for attributes. BI & Reporting Metric Layers Separation of raw measure, business measure, and presentation measure — prevents logic in every dashboard. BI & Reporting Governed Self-Service Self-service on certified datasets and clear contracts — freedom with guardrails instead of shadow IT. BI & Reporting Vanity Metric Metric that looks good but changes no decision — symptom of missing KPI contracts. BI & Reporting North Star Metric One central outcome metric for product/org — needs grain, owner, and supporting metrics. BI & Reporting Leading Indicator Metric that predicts outcomes (pipeline, activation) — complements lagging indicators. BI & Reporting Lagging Indicator Metric that measures outcomes after the fact (revenue, churn) — important, but alone too late to steer. BI & Reporting Metric Proliferation Uncontrolled growth of similar measures — trust drops; certification and deprecation help. Quality Data Contract Agreement between producer and consumer on schema, semantics, SLAs, and breaking-change rules. BI & Reporting Dashboard Sprawl Too many similar dashboards without owner and dataset contract — classic symptom of self-service without guardrails. BI & Reporting Drill-Down Navigate from coarse to fine hierarchy level — needs conformed hierarchies and clear grain. BI & Reporting Drill-Through Jump from aggregate to detail rows or another report — lineage and entitlements must travel with it. BI & Reporting Ambiguous Path Multiple filter paths between tables in the semantic model — yields wrong aggregates if unresolved. BI & Reporting Bidirectional Filter Filters flow both ways on a relationship — powerful and risky for ambiguous paths. BI & Reporting Operational Reporting Day-to-day reporting close to the process — often different latency and grain than analytical marts. BI & Reporting Analytical Reporting Analysis over integrated, historized models — typically mart/semantic layer rather than raw ODS. BI & Reporting Dashboard Consumption surface with visuals, filters, and context for a concrete decision or monitoring task. A dashboard should consume approved metrics; it is not a data catalog, metadata source, or suitable home for competing core logic. BI & Reporting Scorecard Compact KPI overview with targets/status — needs contracts, or you get colorful lights without trust. Quality Fitness for Purpose Quality judged against an explicit use case — not absolute perfection. BI & Reporting OKR Goal system of objectives and measurable key results — key results are often KPIs with owner and grain. BI & Reporting Cohort Group sharing a start trait (signup month) — analyses need stable grain and history. BI & Reporting Attribution Assigning outcomes to touchpoints/causes — highly sensitive to model and definition. BI & Reporting Workspace Collaboration and publish boundary in BI platforms — entitlements, promotion, and certification hang off it. BI & Reporting Dataset Reusable semantic model/package for reports — ideal point for certification and contracts. BI & Reporting Promotion Moving content/datasets from dev/test to prod — needs checks, owner, and rollback. BI & Reporting Excel Export Export from governed surfaces into spreadsheets — often the start of shadow IT and the two-version problem. BI & Reporting Manual Reconciliation Manually reconciling numbers across systems — symptom of missing contracts, grain, and lineage. BI & Reporting Stale Metric Metric with outdated definition or no owner — trust killer despite looking “official”. Quality Data Observability Detect unexpected volume, null, distribution, schema, and freshness anomalies beyond fixed rules. BI & Reporting Conflicting Metric Two measures answer the same question differently — the two-version problem in numbers. BI & Reporting Orphan Report Report without owner or usage — clean up, archive, or deprecate. BI & Reporting Unused Dataset Dataset without consumers — cost and risk without value; usage metadata makes it visible. BI & Reporting Spreadsheet Hell Critical logic and truths only in spreadsheets — end state of shadow IT and Excel export. Quality DQ Gate Hard stop or promote condition in the pipeline based on quality outcomes. BI & Reporting Field Parameter BI control that lets users switch fields/metrics in visuals — governance must constrain allowed dimensions. BI & Reporting Dual Storage Mode A table can use import and DirectQuery together — flexible, but harder to explain and debug. BI & Reporting Incremental Refresh Loads only new/changed partitions instead of a full reload — saves time, needs stable keys and archive policy. BI & Reporting Live Connection The report hits the semantic model live — always current data, but performance and model governance sit centrally. BI & Reporting Import Mode Data is loaded into the BI model and cached there — fast for users, refresh windows and memory become critical. Quality Rule Registry Catalog of DQ checks, owners, severity, and execution context. Quality Remediation Owned fix-and-validate loop for quality and metadata defects. Modeling Staging Area Intermediate zone for raw or lightly transformed data before core model — separates ingest from business logic. Modeling Persistent Staging Staging kept with history — enables rebuilds and audit without re-pulling sources. Modeling Incremental Pipeline Processes only new/changed data since last run — standard for scale, needs clean keys and watermarks. Modeling Ragged Hierarchy Hierarchy with varying depths (levels skipped) — e.g. region without intermediate country node. Modeling Unbalanced Hierarchy Nodes have different child depths — reporting and drill must model parent-child paths explicitly. Modeling Parent-Child Hierarchy Hierarchy via parent ID per row — flexible but queries and BI drill often costlier than snowflake dims. Modeling Hierarchy Path Stored path from root to node — speeds drill and filter in unbalanced hierarchies. Quality Freshness How current data (or metadata] is versus the agreed expectation. Quality Completeness Required fields and records present for the declared purpose. Quality Consistency Same meaning and rules across systems, layers, and replicas. Quality Accuracy Values correctly represent the real-world fact for the use case. Quality Uniqueness No unwanted duplicates at the declared grain or key. Modeling Snapshot Fact Fact table measures state at fixed points in time (e.g. daily inventory) — not every individual transaction. Quality Validity Values conform to allowed formats, domains, and reference sets. Modeling Accumulating Snapshot Fact row follows a process across milestones (order→ship→pay) and is updated — typical for cycle times. Modeling SCD Type 1 Dimension attribute is overwritten — history is lost, suitable for corrections with no analysis need. Modeling Late-Arriving Fact Fact appears after the expected load time — pipelines must handle backfill/replay and dimension keys cleanly. Modeling SCD Type 3 Keeps current and previous attribute values in columns — limited history, not full SCD2. Modeling Galaxy Schema Multiple fact tables share dimensions — common in large DWHs, needs conformed dimensions. Modeling Rapidly Changing Dimension Dimension attribute changes so often that classic SCD2 explodes — often mini-dimension or outrigger instead of full history. Modeling Hard Delete Record is physically removed — irreversible; often needed for erasure, hard on history/audit. Modeling Additive Measure Metric may be summed across dimensions (revenue, quantity) — averages are often not additive. Modeling Semi-Additive Measure Summable across some dimensions, not time (balance, inventory) — grain decides. Modeling Non-Additive Measure Must not be naively summed across rows (ratio, avg price) — aggregation needs a formula, not SUM. Quality Root Cause Analysis (DQ] Trace defect to originating system or process instead of patching marts. Modeling Tombstone Record Marker for deleted entries in logs/streams — consumers must propagate deletes, not ignore them. Modeling Bronze Layer Raw/landing zone with minimal transformation — audit and replay, no business logic. Modeling Silver Layer Cleansed, conformed data — joins and quality, not yet mart-specific. Modeling Gold Layer Business-ready marts/metrics for consumers — grain and KPI definition live here. Modeling Measure Group Group of related measures/facts sharing grain — wrong grouping mixes KPIs. Modeling Domain-Driven Design Model around the business domain (bounded contexts, ubiquitous language) — not the other way around from tools. Quality Quality by Design Bake quality into contracts, models, and pipelines — not “find” it later by sampling dashboards. Quality Anomaly Detection Automatic detection of unexpected patterns in volume, distributions, or metric values — complements rule-based tests. Quality Data Profiling Statistical inventory of columns (nulls, cardinality, patterns) — basis for rules and contracts. Quality Contract as Code Data contracts and assertions versioned in the repo and checked in CI — not only as a wiki paragraph. Quality Timeliness DQ dimension: data arrives in time for the decision — related to freshness, but judged against the use case. Quality Incident Runbook Step-by-step guide for pipeline/DQ incidents — owners, checks, escalation, communication. Quality DataOps DevOps practices for data products: CI/CD, observability, short feedback loops between produce and consume. Quality Quality Score Aggregated score from DQ rules/dimensions — useful as a trend, dangerous as the only truth. Quality Assertion Machine-checkable expectation on data (row counts, ranges, references) — building block of contract-as-code. Quality On-Call Rotating responsibility for off-hours incidents — needs runbooks and clear escalation. Quality MTTR Average time to recover after incidents — runbooks and on-call reduce it. Quality Error Budget Allowed unreliability under the SLO — steers change pace vs. stability. Quality Toil Manual, repetitive ops work without lasting value — automation and golden paths reduce it. Privacy & protection PII Personally identifiable or linkable data. Needs classification, masking, purpose binding, and proven deletion/restriction paths. Quality Orphan Pipeline Pipeline without clear owner/consumer — costs money and creates silent incidents. Privacy & protection DSDR Process and technical ability to execute and evidence data-subject rights (erase/restrict) across systems and lineage. Quality Quarantine Zone Isolated holding area for failed records — pipeline continues, bad data does not reach consumers. Quality Reconciliation Check Compares sums/counts between source and target — finds silent loss that row tests miss. Quality Data Reliability Overall sense that data arrives on time, complete, and trustworthy — more than single DQ checks. Quality Volume Anomaly Unexpected jump or drop in row count — often the first signal of broken upstream jobs. Privacy & protection Retention Rules for how long data may stay active or archived — separated from backup and analytical marts. Privacy & protection Masking Technique to hide or replace sensitive values for unauthorized roles — ideally policy-driven and lineage-aware. Privacy & protection Data Classification Labeling sensitivity or purpose class that drives protection and use. Privacy & protection Sensitivity Label Platform label (e.g. Purview] binding policy to assets or columns. Privacy & protection Purpose Limitation Use only for the agreed purpose — before tooling and mart design. BI & Reporting Report Book Bundled report collection with shared navigation — typical for PDF/print compliance packs. BI & Reporting Mobile BI BI for phone/tablet — layout, offline, and RLS must be designed for small screens. BI & Reporting Embedded Report Report/dashboard embedded in product UI — needs embedding API, auth, and consistent semantic layer. BI & Reporting Report Subscription Email Automated report delivery via email on schedule — distribution layer, not a substitute for interactive BI. BI & Reporting Scheduled Refresh Schedule for data refresh in BI tool — SLA for “fresh” dashboards, often tied to pipeline SLA. Privacy & protection Tokenization Replace sensitive values with reversible tokens under controlled vaulting. Privacy & protection Pseudonymization Reduce identifiability while allowing controlled re-link under safeguards. Privacy & protection Anonymization Irreversible removal of personal identifiability for a stated threat model. Privacy & protection Redaction Drop or blank high-risk fields so they never reach curated or mart layers. Privacy & protection Workforce / Employee Data Policy Separate handling rules for workforce identity vs customer PII in RAW→Mart. Privacy & protection GDPR EU privacy framework with purpose limitation, data-subject rights, and accountability — drives classification, retention, and DSDR. Privacy & protection Legal Hold Freeze against deletion/change due to litigation or investigation — overrides normal retention. Privacy & protection Dynamic Masking Masking at query time by role — raw data stays stored, visibility is controlled. Privacy & protection Hashing One-way transform of identifiers — often for joins without plaintext PII, with collision and rainbow risks. BI & Reporting Aggregation Table Precomputed summary at a coarser grain — speeds reports, must stay consistent with detail source. Privacy & protection Data Minimization Store/share only necessary attributes and rows — tightly linked to purpose limitation and least privilege. Privacy & protection Consent Freely given, informed permission to process — one lawful basis among others, not the only one. BI & Reporting Direct Lake BI reads Parquet/Delta in the lake without classic import — fresh like DirectQuery, often faster than pure warehouse import. Privacy & protection Lawful Basis Legal ground for processing (consent, contract, legitimate interest…) — must fit the purpose — masking is a control, not a lawful basis. BI & Reporting Object-Level Security Hides tables/columns/measures from roles — complements RLS (rows), does not replace it. Privacy & protection Controller Party that determines purposes and means of processing — carries accountability to data subjects. BI & Reporting Drillthrough Navigation from aggregated view to a detail page with context filters — not a substitute for clean grain definition. Privacy & protection Processor Processes personal data on behalf of the controller — needs a contract and evidenced controls. BI & Reporting Tooltip Page Small report page as hover detail — gives context without page navigation, should stay lean. Privacy & protection Special Category Data Especially protected data (health, biometrics…) — stricter requirements than ordinary PII. BI & Reporting What-If Parameter Interactive parameter for scenarios (e.g. price ±10%) — simulates, does not persist facts in the source. BI & Reporting Calculated Column Column materialized in the model (row context) — unlike a measure (filter context); memory-heavy on large tables. Privacy & protection Synthetic Data Artificially generated data for test/training — reduces PII risk, needs realism checks. Privacy & protection Differential Privacy Mathematical privacy protection via controlled noise — strong, but utility trade-off. BI & Reporting Visual-Level Filter Filter applies to one visual only — not the whole page/report; easy to miss when debugging. Privacy & protection Data Residency Requirement for where data may physically/legally reside — drives cloud region and sharing design. BI & Reporting Page-Level Filter Filter for all visuals on a report page — scope between visual- and report-level. Privacy & protection PII Scan Automated search for personal patterns in schemas/content — starting point for classification. BI & Reporting Report-Level Filter Filter applies report-wide across pages — powerful, but risky when users do not see the scope. BI & Reporting Sync Slicer Slicer selection stays synced across pages — less clicking, more coupling. Privacy & protection Tombstone Marker that a record is deleted/suppressed — relevant for CDC, DSDR, and soft deletes. Privacy & protection k-Anonymity Each quasi-identifier profile appears at least k times — classic, limited anonymization approach. BI & Reporting Paginated Report Print/PDF-oriented report with fixed layout — unlike interactive dashboards. Privacy & protection Format-Preserving Encryption Encryption that preserves format/length — for legacy fields; does not replace access policies. BI & Reporting On-Premises Data Gateway Bridge from cloud BI to on-prem sources — needs HA, permissions, and monitoring like production infra. BI & Reporting Bookmark State Saved filters/views in reports — personal vs shared must be clear or “wrong truth” spreads. Privacy & protection Homomorphic Encryption Compute on encrypted data without decrypting — powerful, still often costly/complex today. Privacy & protection Encryption at Rest Encryption of stored data — baseline hygiene; does not protect against authorized queries. BI & Reporting Report Theme Central colors/fonts for reports — without a theme, every dashboard becomes a design experiment. Privacy & protection Encryption in Transit Encryption on the wire — standard for all pipeline and API paths. BI & Reporting Mobile Report Report layout for small screens — different priorities than desktop, not just scaled dashboard. Privacy & protection Health Data Health-related personal data — usually special category / especially protected. BI & Reporting Report Subscription Automated delivery of report snapshots — snapshot time and filters must be documented. Privacy & protection Employee Data Personal employee data — own policies, co-determination, and purpose limits. BI & Reporting Visual Cross-Filter Click in one visual filters others — powerful for exploration, easy to miss when debugging. BI & Reporting Button Slicer Filter as clickable buttons/tiles — UX-friendly, but limited cardinality recommended. BI & Reporting Relative Date Filter Filter like “last 7 days” relative to today — dynamic; timezone must be correct. Privacy & protection Data Clean Room Controlled environment for joint analytics without raw data exchange — queries yes, copying datasets no. BI & Reporting Top N Filter Shows only top/bottom N by measure — tie-breaker and filter context must be documented. Privacy & protection Retention Schedule Defines how long each data class is kept and when it is deleted/archived — without a plan compliance risk grows. BI & Reporting Small Multiples Same chart repeated per category — compares patterns, not absolute scale alone. BI & Reporting Bookmark Navigation Buttons jump to saved filter/page states — guided analytics instead of free exploration. Privacy & protection Right to Erasure Right to have personal data erased — needs lineage through backups and downstream copies. Access & security RBAC Access by role membership. Access & security ABAC Access by attributes (clearance, purpose, residency, etc.]. Access & security Least Privilege Minimum access needed for the job — default deny elsewhere. Access & security Segregation of Duties (SoD] Split conflicting duties (e.g. grant vs approve] to reduce abuse risk. Access & security Section Access Qlik row-reduction security model in apps. Access & security Row-Level Security (RLS] Filter rows by user or role claims in warehouse or BI. Access & security Access Recertification Periodic re-approval that entitlements are still needed. Quality Distribution Check Checks whether value/category distribution looks expected — early drift alarm before hard business-rule fails. Quality Uniqueness Constraint Rule that keys or combinations must be unique — basis for referential integrity and aggregation. Quality Referential Check Validates foreign keys point to existing parent rows — prevents orphaned facts in star schemas. Quality Schema Validation Checks columns, types, and required fields against expected schema — often gate before lakehouse write. Quality Anomaly Detection Rule Automated rule for unusual values or counts — complements static thresholds on volatile metrics. Quality Profiling Job Scheduled job for null rates, distinct counts, and patterns — input for quality rules and anomaly detection. Quality Quality Gate Checkpoint in CI/CD/pipeline — deploy or promote only when defined DQ checks pass. Quality Observability Alert Alert from metrics/logs (latency, freshness, error rate) — connects ops visibility with data quality. Access & security IAM Identity and access management backbone for authenticating subjects to data systems. Access & security Column-Level Security Access control at column grain — complements RLS when entire attributes must be invisible to roles. Access & security Encryption Protection of data in transit and at rest — baseline hygiene; replaces neither masking nor RBAC. Access & security Entitlement Concrete access right of an identity to an asset — subject of recertification and least privilege. Access & security Service Principal Non-human identity for pipelines and apps — own entitlements, rotation, and recertification. Access & security Audit Log Evidence of who accessed or changed what when — basis for recertification and incidents. Access & security SSO One login for many apps — simplifies IAM, but also centralizes risk. Access & security MFA Multiple authentication factors — baseline for privileged and PII access. Access & security SCIM Standard for automated provisioning/deprovisioning of identities into apps. Access & security Zero Trust No implicit trust from network zone — continuous verification of identity and context. Access & security Secrets Management Central, rotatable store for keys/tokens — no secrets in repos or notebooks. Access & security Key Vault Managed service for keys and secrets — often the anchor for encryption and BYOK. Access & security BYOK Customer controls encryption keys — increases control and compliance requirements. Access & security Row Access Policy Policy object controlling row access in the warehouse — complements app-side RLS. Access & security Object Tag Metadata tag on tables/columns for classification and policy binding — active metadata in action. Access & security Just-in-Time Access Time-bound rights only when needed — reduces standing privileges. Access & security PAM Management of highly privileged access — vaulting, session control, recording. Access & security Break-Glass Access Emergency access with strong audit — exception, not a steady state. Access & security Column Masking Policy Policy that role-masks column values — often bound to tags/classification. Access & security Tag-Based Access Access via classification/object tags instead of object names alone — scales better in large estates. Metadata Lineage Traceable origin and transformation of data — from source through pipelines to reports and deletion paths. Metadata Data Catalog Search and relationship surface for data assets, terms, owners, policies, and consumers. A data catalog makes metadata discoverable and connected; it is not automatically the system of record for every metadata item and not a dashboard or semantic layer. Quality Freshness SLA Promised maximum data age until use — without measurement, “fresh” is marketing. Metadata Metadata Descriptive and controlling facts about data, reports, models, or processes: schema, meaning, owners, classification, lineage, quality, policy status, and time context. Metadata is the governance content; the catalog is only one possible surface for it. Quality Null Rate Share of missing values in a column/time window — early warning for broken feeds and schema drift. Quality Data Diff Compares two datasets/versions at row or aggregate level — finds drift after refactors and migrations. Quality Expectation Suite Versioned set of declarative quality rules (range, uniqueness, references) — tests as code, not ticket prose. Quality Completeness Check Checks whether expected fields/rows are present — “correct but incomplete” still fails. Metadata Metadata Catalog Catalog view that indexes metadata from different sources and makes it usable. The term usually describes the surface or platform, not automatically the business authority for every definition, classification, or approval. Quality Uniqueness Check Ensures keys/combinations are unique — duplicates break joins and KPIs. Metadata Control Plane Control layer that manages rules, policies, identities, approvals, lineage, or operating states. It may prepare decisions or enforce them technically, but it does not replace the accountability of owner, steward, or custodian. Quality Validity Rule Value must fall in an allowed domain (range, enum, format) — “not null” alone is not enough. Quality Duplicate Rate Share of duplicate keys/rows in a window — early signal for broken upserts and CDC gaps. Quality Orphan Record Row without a valid parent/dimension reference — joins yield gaps or wrong “unknown” buckets. Quality Accuracy Check Checks whether values are factually correct (reconcile with source/golden set) — technically “green” is not enough. Metadata Business Glossary Agreed business terms and definitions — related to, not identical with, the catalog. Quality Consistency Check Same metric/entity aligns across systems — DWH vs CRM contradictions are classic. Quality Conformity Check Data matches format/standard (ISO date, currency code) — schema ok, content still unusable. Quality Integrity Check Relationships between entities hold (FK, cardinality) — orphan rows are an integrity fail. Quality Shift-Left Quality Quality rules and tests as early as pipeline/dev — avoid expensive fixes in the dashboard. Quality Timeliness Check Checks whether data arrives on time — distinct from freshness SLA (age since last load). Quality Volume Check Monitors expected row volume per load — sudden halving is often worse than a schema error. Quality Range Check Values must fall in allowed interval — catches unit errors and overflow early. Quality Data Quality Scorecard Aggregated view of DQ dimensions per product — score without action plan is decoration. Quality Quality Dimension Axis of data quality (accuracy, completeness, timeliness …) — shared vocabulary for scorecards. Metadata Active Metadata Metadata that drives automation (policies, quality, routing], not only documentation. Metadata Metadata Provenance Who or what authored a metadata claim and from which source of truth. Metadata Metadata Harvesting Automated collection of technical metadata from platforms into the catalog. Metadata Impact Analysis Trace downstream blast radius of a field, model, or KPI change via lineage. Metadata Column Lineage Field-level origin and transform path (needed for PII propagation and DSDR]. Metadata Metadata Enrichment Add business context, owners, classification, and KPI links to technical assets. Metadata AI-Ready Metadata Complete, current, permitted-use metadata suitable for assistants and RAG. Metadata dbt meta Structured metadata in dbt YAML used to drive governance automation. Metadata Centralized vs Federated Metadata Decide per capability what is central discovery vs domain-authored truth. Privacy & protection Privacy by Design Build privacy into architecture and process from the start — not as a last filter before go-live. Privacy & protection Data Masking Obscures sensitive values in non-prod/self-service (hash, partial, fake) — not a substitute for prod access control. Privacy & protection DPIA Structured risk analysis before high-risk processing — documents mitigations, does not replace ongoing controls. Privacy & protection Consent Management Captures, versions, and propagates consent in machine-readable form — a marketing opt-in alone is not proof. Privacy & protection Controller vs Processor Controller decides purposes/means; processor acts on instructions — mixing them up costs contracts and liability. Privacy & protection Legitimate Interest Legal basis after balancing interest vs data-subject rights — not a free pass without documentation. Metadata OpenLineage Open standard for job and dataset lineage events — interoperable across orchestrators and catalogs. Privacy & protection Data Subject Natural person to whom personal data relate — holder of access and erasure rights. Privacy & protection Cross-Border Transfer Transfer of personal data to other jurisdictions — needs a transfer mechanism and risk analysis. Privacy & protection Schrems II CJEU ruling on US transfers: Privacy Shield invalid — transfer impact assessments and supplementary measures became central. Privacy & protection Retention Policy Rules for how long which data is kept for what purpose — without technical enforcement, policy is folklore. Privacy & protection Data Breach Unauthorized disclosure/access to personal data — notification deadlines and forensics are mandatory. Metadata Technical Metadata Schema, types, jobs, storage settings — what systems know about data structures and pipelines. Privacy & protection Breach Notification Duty to inform authority/subjects about relevant incidents — needs playbook and evidence. Privacy & protection Re-Identification Risk Risk of linking pseudonymized/anonymized data back to a person — watch quasi-identifiers. Privacy & protection Standard Contractual Clauses Contract modules for third-country transfers — complement TIA and technical measures, do not replace them. Privacy & protection Data Processing Agreement Contract between controller and processor on purpose, subprocessors, deletion — cloud without DPA is a red flag. Privacy & protection Privacy by Default Most restrictive privacy setting is default — users must actively expand, not fight opt-out. Privacy & protection Transfer Impact Assessment Risk analysis for third-country transfers post-Schrems II — complements SCCs, does not replace them. Privacy & protection Sub-Processor Processor engages another vendor — chain must be transparent and controllable in the DPA. Privacy & protection Consent Record Proof of who consented to what and when — must propagate to downstream systems. Metadata Business Metadata Business meaning, owners, glossary terms, and usage hints — makes technical assets understandable. Metadata Operational Metadata Runtimes, status, incidents, SLAs — metadata from operating pipelines and jobs. Metadata Usage Metadata Who uses which assets how often — basis for prioritization, certification, and cleanup. Privacy & protection Data Subject Request Individual request for access, deletion, or correction — must be traceable across systems and lineage. Privacy & protection DSR Workflow Standard process from ticket to deletion/export with deadlines — without workflow, copies linger in backups. Privacy & protection Opt-out Objection to use (marketing, profiling) — must be reflected technically in segments and pipelines. Privacy & protection Opt-in User actively consents — standard for many marketing and analytics use cases in the EU. Privacy & protection Legitimate Interest Assessment Documented balancing for processing without consent — needs purpose, necessity, and balancing. Metadata Declared Metadata Explicitly documented or code-declared metadata (e.g. dbt meta, glossary) — intent, not observation. Metadata Detected Metadata Harvested or inferred metadata from systems — schema, lineage, usage — complements declared metadata. Metadata Control-Driving Metadata Metadata that actively drives policies, masking, routing, or gates — not merely describes. Process RACI Role matrix for decisions: who executes (R], who owns (A], who is consulted (C], who is informed (I]. Metadata Metadata Graph Connected view of assets, owners, lineage, and policies — foundation for impact analysis and active metadata. Metadata Business Lineage Lineage in business terms (KPI ← mart ← domain) — understandable for owners and stewards, not only engineers. Metadata Technical Lineage Lineage at table/column/job grain from harvesting and runtime — precise for impact and debugging. Metadata Enterprise Vocabulary Shared business language across domains — mapping local labels to canonical glossary terms. Process KPI Governance Clear definition, owner, and change process for metrics — prevents conflicting numbers across tools and meetings. Metadata Metadata Product Metadata with product thinking: owner, SLA, versioning, and consumers — not just a catalog dump. Metadata Schema Registry Central management and evolution of event/message schemas — compatibility checks before breaking changes. Metadata Event-Driven Metadata Metadata updates as events from jobs/catalogs — instead of only nightly full harvests. Metadata Static Metadata Slow-changing metadata (owner, glossary, classification) — often declared, not harvested every minute. Metadata Dynamic Metadata Frequently changing metadata (freshness, usage, job status) — typically detected/observed from runtime. Process Decision Rights Who may decide purpose, access, definitions, and exceptions — and at what risk tier. Access & security Privileged Access Management Controls and monitors highly privileged accounts (admin, break-glass) — just-in-time instead of standing admin. Access & security Service Account Non-human identity for pipelines/apps — needs rotation, scope, and an owner, not a shared wiki password. Metadata Source-Local Metadata Metadata that originates and is maintained in the source tool — harvesting pulls it into the graph. Access & security Break-Glass Access Time-boxed emergency access with audit — an exception reviewed afterward, not an everyday path. Metadata Descriptive Metadata Metadata that explains and helps discovery — unlike control-driving metadata that steers systems. Access & security Workload Identity Identity for workloads (pods, jobs) without long-lived secrets — short-lived tokens instead of passwords in repos. Access & security Conditional Access Access depends on signals (device, location, risk) — MFA alone is often too coarse. Metadata Distributed Metadata Metadata across many tools/domains without a mandatory single catalog — needs federation and a clear source of truth per aspect. Access & security Secrets Rotation Regular replacement of keys/passwords — without automation, secrets stay “valid forever”. Metadata Bounded Context Bounded meaning space for terms and models — prevents “customer” meaning the same thing everywhere. Access & security Identity Provider Central authority for authentication/SSO — apps trust tokens instead of their own password DBs. Access & security OAuth Client Registered app with client ID/secret and scopes — rotation and least-privilege scopes are mandatory. Access & security API Gateway Central entry for APIs (auth, rate limit, routing) — policies here instead of duplicating in every service. Access & security Mutual TLS Client and server authenticate via certificates — stronger than TLS-only; needs rotation and CA governance. Metadata Data Dictionary Technical catalog of tables/columns/types — complements the business glossary, does not replace it. Access & security Network Segmentation Split network into zones to hinder lateral movement — keep data plane and admin separate. Process Operating Model Cadence, handoffs, capacity, and escalation that make roles real. Metadata Passive Metadata Automatically harvested schemas/stats without active curation — good for discovery, weak for binding semantics. Access & security Security Posture Overall picture of controls and gaps — score alone is not enough without prioritized remediation. Access & security Threat Modeling Systematic analysis of threats and mitigations — before go-live, not after the incident. Access & security Attack Surface Sum of exposed entry points (APIs, accounts, data exports) — shrinking beats more alerts. Access & security Public Key Infrastructure Infrastructure for certificates and trust chains — foundation for mTLS and signed tokens. Access & security Phishing-Resistant MFA MFA that resists replay/phishing (e.g. FIDO2) — SMS codes do not count. Process Escalation Path Defined route when Steward, Owner, or Platform cannot resolve within SLA. Process Stewardship Intake Prioritized entry path for definition, DQ, and classification work. Process Stewardship Capacity Model FTE or protected-time model so stewardship is funded, not „on the side“. Process Governance Cadence Recurring forums and reviews (classification, access, KPI, council]. Process Data Lifecycle Create → use → retain/archive → delete/retire with accountable stages. Process Role Sprawl Too many overlapping RACI hats that dilute accountability. Process Federated Governance Central standards plus domain execution (vs pure central or pure local]. Process Policy as Code Enforceable access, quality, and privacy rules in versioned, testable artifacts. Process Data Technical Debt Accumulated compromises in pipelines, models, and reports — interest paid as incidents, shadow IT, and slow change. Process Accountability Who finally owns outcome and risk — the “A” in RACI, not the same as doing the work (“R”). Process KPI Operating Model Cadence, roles, and change process around KPIs — from definition through certification to breaking changes. Process ADR Short record of an architecture decision including context and consequences — prevents tribal knowledge. Process Tribal Knowledge Critical knowledge only in a few heads — bus-factor risk; metadata and ADRs are antidotes. Process GDR Decision record of a governance decision: the decision, the A, scope and kill line — not meeting minutes. Process Two-Version Problem Two “official” numbers for the same question — typical without a metric store and decision rights. Process SRE Practice of running reliability with engineering — SLOs, error budgets, less toil. Process Platform Engineering Builds self-service platform products for teams — golden paths instead of ticket-driven ops. Process Golden Path Supported standard path (templates, CI, patterns) — fast and compliant at once. Process Value Stream End-to-end flow from need to outcome — helps prioritize data products and roadmaps. Process Lead Time Time from request to production-ready — core flow metric beside cycle time. Process Cycle Time Time of active work on an item — drops with smaller batches and less WIP. Process WIP Work open at the same time — high WIP lengthens lead time and creates half-done work without outcome. Process Product Thinking Focus on user outcomes, lifecycle, and ownership — instead of pure ticket/report production. Process Community of Practice Network for sharing standards and learning across teams — complements a CoE, does not replace it. Process Bus Factor How many people can be hit by a bus before knowledge/system fails — low with tribal knowledge. Process Ticket-Driven Delivery Work only via ticket queues — typical opposite of product thinking and golden paths. Process Product-Driven Delivery Delivery via products, outcomes, and roadmaps — instead of endless ad-hoc tickets. Process Hero Culture Systems run only via individual heroics — bus factor and burnout instead of reliability. Process Hub and Spoke Central hub plus domain spokes — org pattern for governance and platform. Process Inner Source Open-source practices internally — PRs, shared ownership, visible standards. Process Trunk-Based Development Short branches, frequent merges to trunk — fits CI/CD and small batches. Process Canary Release New version first for a small traffic share — limits blast radius; monitoring decides further rollout. Process Shadow Deployment New pipeline/model runs in parallel without user impact — compares results, switches later. Process Chaos Engineering Deliberately injecting failures in production/staging to prove resilience — not random breakage. Process Circuit Breaker Temporarily stops calls to an unhealthy downstream — prevents cascades, needs clear recovery rules. Process Value Stream Mapping Visualizes end-to-end flow from idea to value — makes queues and handoffs in data/BI visible. Process Blue-Green Deployment Two parallel environments; traffic switches atomically — faster rollback than patching live step by step. Process Blameless Postmortem After incidents, analyze causes and system gaps without blame — focus on fixes and learning. Process Runbook Concrete steps for ops/incident (checks, rollback, contacts) — a living document, not a one-off wiki page. Process WIP Limit Cap on parallel work — protects throughput and quality, including in data/BI backlogs. Process Definition of Done Shared done criteria (tests, docs, owner, monitoring) — “done” without a DoD is only locally done. Process Definition of Ready Criteria for when work is start-ready (scope, owner, data availability) — prevents half-baked starts. Process Incident Severity Incident severity rating (Sev1…n) drives response time and command structure — without it, everything is “urgent”. Process Incident Commander Coordinates response, communication, and decisions in an incident — not necessarily the deepest debugger. Process Error Budget Policy Defines what happens when the error budget is spent (freeze, reliability focus) — policy without consequence is theater. Process Pair Programming Two people on one problem (driver/navigator) — knowledge transfer and early review, not “twice the cost”. Process Continuous Delivery Keep software/data pipelines releasable at any time — manual “manual end-of-week deploy” is an anti-pattern. Process Change Advisory Board Forum for risky changes (prod, schema) — should reduce gatekeeping, not delay every ticket. Process RTO Maximum acceptable downtime until recovery — without RTO, DR is PowerPoint only. Process RPO Maximum tolerable data loss in time — backup frequency must match. Process Ubiquitous Language Shared, precise terms in team and code — prevents “revenue” synonyms in SQL and slides. Process Context Map Visualizes relationships between bounded contexts (upstream/downstream) — clarifies integration and ownership. Process Trunk-Based Development Short-lived branches, frequent merge to main — fits CI/CD and DataOps pipelines. Process Game Day Planned stress test/incident exercise in prod-like env — finds gaps before real outages. Process On-Call Rotation Schedule of who carries alerts/incidents when — without rotation, one hero burns out and knowledge silos. Access & security Security Event Logged security-relevant action (login, policy change) — raw material for SIEM and forensics. Access & security SIEM Alert Alert from SIEM correlation of multiple events — quality depends on parsing, tuning, and false-positive handling. Access & security SOC Analyst Triages alerts, escalates incidents, documents playbooks — operational heart of security operations. Access & security Vulnerability Scan Automated search for known weaknesses in images/hosts — findings need prioritization and patch SLA. Access & security Penetration Test Simulated attack on systems with report — deeper than scan but point-in-time, not continuous monitoring. AI & ML RAG (Retrieval-Augmented Generation] Ground LLM answers in retrieved governed documents or data. AI & ML AI Guardrails Controls that constrain prompts, tools, and outputs for safety and compliance. Metadata Column Profiling Statistics about columns (nulls, cardinality, patterns) — basis for DQ and catalog enrichment. AI & ML Prompt Injection Attack that hijacks model behavior via malicious content in inputs or context. Metadata Data Lineage Graph Graph from sources→transforms→consumers — impact analysis needs edges, not wiki prose alone. Metadata Open Metadata Open platform for catalog, lineage, and governance workflows — metadata as a product, not a side-car. Metadata Metadata Lineage Metadata about origin and transformation — distinct from business-glossary prose alone. Metadata Business Glossary Term Business-defined term with owner and status — links to technical assets, does not replace them. Metadata Semantic Metadata Metadata about meaning (terms, KPIs, synonyms) — links glossary to technical assets. Metadata Catalog Ingestion Automatic loading of technical metadata into the catalog — without enrichment, just inventory. AI & ML Hallucination Confident model output not grounded in retrieved or training evidence. AI & ML Human-in-the-Loop Mandatory human review or approval for high-risk AI actions. AI & ML Training Data Datasets used to train or fine-tune models — need lineage, rights, and quality. AI & ML Machine Unlearning Methods to reduce or remove the influence of specific training points on a model — usually approximate in practice and paired with verify; often an alternative or complement to retrain/adapter drop and index rebuild. AI & ML AI Ingestion Gate Control point before retrieval index or training ingest: hygiene disposition, PII/secret rate, purpose approval, and metadata contract must pass — otherwise block or a time-boxed waiver. AI & ML Shadow AI AI use outside sanctioned paths — e.g. public chats, personal API keys, uncontrolled vendor POCs — without owner, purpose card, and evidence. AI & ML Sanctioned AI Path Catalogued AI usage path with purpose, allowed data classes, owner/steward, risk class, logging, and exit criteria — prerequisite for production operation. AI & ML Corpus Poisoning Intentional or negligent contamination of retrieval/training corpora (e.g. prompt-injected docs, tainted vendor packs) — needs trust tiers and integrity gates, not hygiene alone. AI & ML Inference Runtime model execution against new inputs. AI & ML Feature Store Governed reuse of ML features with versioning and serving contracts. AI & ML LLM Large language model for text/code — needs guardrails, grounding, and clear rights on context sources. AI & ML Embedding Vector representation of text or features for similarity search — foundation for RAG and semantic search. AI & ML Vector Store Storage and index for embeddings with similarity retrieval — often coupled to RAG pipelines. AI & ML AI Agent LLM-driven unit that plans and executes tools/steps — needs guardrails and human-in-the-loop. AI & ML Fine-Tuning Further training a base model on your data — don’t forget rights, lineage, and evaluation. AI & ML Model Registry Versioned store of ML models with metadata, stages, and approvals — analogous to a data product catalog. AI & ML MLOps Operating and delivering ML models: CI/CD, monitoring, retraining, and governance along the model lifecycle. AI & ML AI Evaluation Systematic measurement of quality, safety, and regressions for models/agents — before and after deploy. AI & ML Prompt Engineering Crafting prompts and system instructions — replaces neither grounding nor guardrails and rights checks. Metadata Business Term Link Link business term ↔ technical field/asset — makes catalog understandable for non-engineers. AI & ML Grounding Bind model answers to retrieved/verified evidence — core of RAG against hallucination. AI & ML Chunking Splitting documents into retrieval units — chunk size steers recall and noise. Metadata Technical Term Metadata entry for column, topic, or API field — often with type, lineage, and owner in catalog. Metadata Catalog Search Full-text/semantic search across datasets, terms, and dashboards — quality depends on tags and business links. AI & ML Vector Search Similarity search over embeddings — basis for RAG and discovery in metadata/content. AI & ML Context Window Max tokens a model sees at once — limits prompt, history, and retrieved chunks. Metadata Metadata API API for catalog metadata (assets, lineage, tags) — basis for automation and IDE integration. Metadata Lineage Export Export of lineage graph for audit, impact analysis, or external tool — not just UI screenshot. AI & ML Tool Calling LLM selects and calls tools/APIs — needs guardrails, auth, and observability. Metadata Impact Report Report of which reports/pipelines are affected by schema or term change — from lineage/catalog. AI & ML Model Card Standardized docs on purpose, data, limits, and risks of a model — governance artifact. AI & ML Data Drift Shift in input distributions vs. training — needs monitoring and retraining. AI & ML Concept Drift Change in the input→target relationship — model performance drops despite “same” features. AI & ML LoRA Parameter-efficient fine-tuning — smaller adapters instead of full retrain. AI & ML Structured Output Model answer in a fixed schema (JSON etc.) — eases automation and validation. Process DMBOK Data Management Body of Knowledge (DAMA): shared professional canon for governance, metadata, quality, security, and lifecycle — at Binom bridged via the 8 pillars, not a chapter-by-chapter clone. AI & ML Model Context Protocol Open protocol so assistants connect tools/data sources in a standard way — governance must whitelist allowed tools. AI & ML Agentic Workflow An LLM plans and executes steps with tools — needs guardrails, observability, and clear stop criteria. AI & ML Reranker Second stage after retrieval: re-scores candidates for relevance — costlier than pure vector search, often more precise. AI & ML Chain of Thought Prompt technique where the model emits intermediate steps — can improve reasoning, but leaks thinking traces. AI & ML RLHF Fine-tuning with human feedback/reward — steers behavior, does not replace domain guardrails. AI & ML Few-Shot Prompting Prompt includes a few examples of the desired format — steers output without fine-tuning, consumes context window. AI & ML Jailbreak Attempt to bypass a model’s safety/policy boundaries — guardrails and monitoring must resist it. AI & ML Vector Database Stores embeddings and searches by similarity — foundation for RAG, not a substitute for relational truth. AI & ML Temperature Controls randomness of token sampling — low = more deterministic, high = more creative/error-prone. AI & ML AI Act EU regulation with risk-based duties for AI systems — governance and documentation become mandatory, not optional. Process CDMP DAMA International certification for DMBOK body knowledge — shared professional language; practical start at Binom via pillars, hub, and tools. AI & ML Multi-Agent System Multiple specialized agents coordinate tasks — needs clear roles, shared state, and stop rules. AI & ML Eval Harness Repeatable test battery for LLM/agent quality (golden sets, scorers) — without evals, “better” is only a feeling. AI & ML Red Teaming Deliberately attacking models/agents (jailbreak, injection, bias) — finds gaps before production users do. AI & ML Zero-Shot Prompting Task with no examples in the prompt — cheap on context, often weaker than few-shot for format tasks. AI & ML System Prompt High-priority steering instruction (role, rules) — should be versioned and reviewed like code. AI & ML Prompt Template Reusable prompt structure with placeholders — maintain centrally instead of copy-paste across ten services. AI & ML Semantic Caching Caches answers for semantically similar queries — saves tokens, risks stale/wrong hits without TTL/policy. AI & ML Groundedness Claims are supported by provided sources — a metric against hallucinations in RAG. AI & ML Model Routing Chooses model/route per request (small/large, specialized) by cost, latency, and risk — governance must whitelist routes. AI & ML AI Observability Traces, metrics, and evals for prompts/agents (latency, cost, quality, safety) — without telemetry, no operations. AI & ML Agent Memory Persistent context for agents across sessions — needs retention, isolation, and deletion paths. Process DCAM Maturity/assessment model for data management programs — measures capability; Binom pillars and advisor help find gaps and artifacts, but do not replace a DCAM assessment. AI & ML Speculative Decoding Small model proposes tokens, large model verifies — cuts latency, needs compatible model pairs. AI & ML Prompt Chaining Output of one prompt step feeds the next — needs clear interfaces and per-step evals. AI & ML LLM Gateway Central proxy for LLM calls (auth, routing, logging, limits) — governance point before the model. AI & ML Chunking Strategy How documents are split for RAG (size, overlap, structure) — bad chunking kills retrieval. AI & ML Hybrid Search Combines keyword/BM25 and vector search — better for exact IDs and semantic paraphrases. AI & ML Context Compression Shrinks context before the model (summaries, top-k) — saves tokens, risks information loss. AI & ML Embedding Model Model that maps text to vectors — model choice affects retrieval quality and cost. AI & ML Cross-Encoder Scores query-document pairs jointly — more precise than bi-encoder, costlier at large candidate sets. AI & ML Tool Schema Machine-readable description of allowed tool parameters — without schema, models hallucinate arguments. AI & ML Agent Planner Component that plans steps/tools before execution — without planner, agent becomes ReAct roulette. Process Artifact Reusable, versionable output of governance or pipeline work — e.g. schema, policy draft, incident record, or catalog form; semantics stay stable even when the tool changes. AI & ML Agent Executor Runs planned tool calls and collects results — needs timeouts, retries, and audit log. AI & ML Output Moderation Filters/classifiers on model output (toxicity, PII, policy) — input guard alone is not enough. AI & ML Human Feedback Loop Humans rate/correct model outputs for training or live escalation — needs clear rubrics. Architecture Content Estate Inventory and operating boundary of all collaboration/content systems (mail, sites, files, chat) before they feed AI or analytics. Process Evidence Pack Repeatable pack of policies, exports, and contextual screenshots with owners — for audit, exit, or incident, not a loose screenshot archive. Process Scope Drift Gradual expansion of purpose, data scope, or use beyond contract/review — common with partners and AI pipelines. Process Business Function Organizational function (sales, finance, HR, IT, engineering, ops, marketing, partner) as a domain entry for pain points, controls, and matching stories — complements roles, does not replace them. Process Knowledge Check Hub for series quizzes, glossary buzzword quiz, and learning progress — checks understanding from stories and terms without a separate LMS. Process Story Quiz Embedded quiz block in story Markdown (quiz fence) — checks a story or series core points and stores progress locally. AI & ML RAG Eval Operational assessment of retrieval vs answer quality (relevance, grounding, hallucination) — distinct from pure model eval. AI & ML Sanctioned AI Officially approved AI use with owner, policy, eval, and HITL — counterpart to shadow AI. AI & ML AI Acceptable Use (AUP) Path-bound rules for which AI use is allowed or forbidden, including escapes and escalation — bound to sanctioned-paths-v1, not a slogan PDF. AI & ML AI Impact Pack (FRIA) Operating artifact for impact/fundamental-rights assessment before high/limited release — a gate, not just a document pile. AI & ML AI Red-Team Pack Recurring adversarial test pack (jailbreak, prompt injection, tool abuse) as a pre-prod gate for sanctioned paths. AI & ML Rights Register (AI) Register that separates license/TDM/copyright flags for retrieval, training, and redistribution — catalog gates instead of implicit assumption. AI & ML Embedded AI Surface AI capability embedded in IDE, BI, CRM, mail, or tickets — needs inventory, scope, and scorecard like a sanctioned path. AI & ML GPAI Deployer Role that puts an AI system into service or uses it — distinct from the GPAI provider; deployer duties are not replaced by provider docs. Metadata Ontology Agreed meaning model (entities, relations, terms) — needs a sync contract when catalog and platform ontology diverge. BI & Reporting KPI Measure of goal attainment — needs owner, definition, grain, and time logic, or it becomes a vanity metric. Privacy & protection DSAR Data-subject request for access or other rights — needs discoverable systems and auditable deadlines. Privacy & protection ROPA Record of processing activities — purposes, categories, recipients, and retention must stay current and auditable. Process DORA EU regulation on digital operational resilience for finance — ICT risk, testing, and third-party oversight. Access & security NIS2 EU directive on cybersecurity and incident reporting for essential and important entities. AI & ML GPAI General-purpose AI models with broad applicability — provider vs deployer duties differ under the EU AI Act. Process Mob Programming Whole team works on one problem (one driver) — strong for complex pipeline/model topics. Process Incident Retrospective Structured review after incident — focus on system improvement, not blame assignment. Process Sprint Review Demo and feedback at sprint end — for data teams often “show new dataset/metric”, not just tickets. Process Capacity Forecast Forecast of available team capacity vs. backlog — prevents overpromising on data roadmap. Process Toil Reduction Automating repetitive ops (manual reruns, ticket-driven fixes) — SRE principle for data platform teams. AI & ML High-Risk AI AI systems with elevated risk under the AI Act — requiring risk management, data governance, and documentation. AI & ML Fairness Assessing and controlling unequal model impacts across groups — criteria and thresholds must be defined upfront. AI & ML Adverse Impact Demonstrable disadvantage to a group from a decision or model — a gate before production use. AI & ML Bias Systematic skew in data or model behavior — requires measurement, mitigation, and documented residual risk. AI & ML Explainability Traceability of model outputs for decision-makers and affected people — method must match the risk class. BI & Reporting MQL Marketing-qualified lead by agreed criteria — definition and sales handoff must be contractually clear. BI & Reporting LTV Expected net value of a customer over the relationship — grain, discounts, and churn assumptions must be documented. BI & Reporting Churn Loss of customers or revenue in a period — numerator, denominator, and reactivation logic must be unambiguous. BI & Reporting Forecast Forward-looking estimate of metrics — version, assumptions, and separation from budget must be versioned. BI & Reporting Budget Approved plan value for a period and org unit — must not be confused with forecast or actuals. Data & products Ledger Accounting system of record for actuals — management accounting and analytics must respect the boundary. Data & products Headcount Count of people or FTE — as-of date, employment status, and org assignment must be defined. AI & ML Foundation Model Large pretrained model used as a base for many downstream apps — often GPAI under the AI Act. AI & ML Citation Traceable source attribution for generated claims — a gate for BI copilots and compliance-sensitive answers. AI & ML Copilot Embedded AI assistant in productivity tools — needs allowed paths, context bounds, and output gates. AI & ML Retrieval Pipeline Chain index → retrieve → rerank → context for LLM — quality and latency depend on each step. AI & ML Vector Chunk Text segment plus embedding in vector store — chunk size affects recall and hallucination risk. AI & ML Semantic Search Index Index for similarity search on embeddings — core of catalog search and enterprise RAG. AI & ML Query Expansion Broadens user query (synonyms, LLM) before retrieval — helps with domain-specific terms. AI & ML Function Calling API LLM API layer for structured tool calls — basis for agents with database and API actions. AI & ML Index Rebuild Full rebuild of search/vector index — needed on embedding model change or chunk policy shift. BI & Reporting Threshold Alert A message when a threshold is crossed — with an owner and a next step. Not a mini-dashboard in the mail and not a substitute for the close dashboard. BI & Reporting Data Export A file or feed for a named purpose (close, audit, upstream system). Not an interactive dashboard and not a chart collection as consolation. Process Corridor (follow-on work) What domain A leaves so domain B can join: keys, grain, artifact names, and one A per decision type — not a central glossary. Process Join Pack Operating artifact `join-pack.md`: freeze, leave-open, brief-customer, neighbor countersign, and join test. Practice after the first review — not Sales. Process Freeze / leave-open What must hold for every follow-on (keys, grain, filenames, one A) versus what stays local or deliberately undecided (semantic layer, harvest, “one number”). Process Project moment The week in the customer project (CRM go-live, close, mart cut) where the governance SKU attaches — not the Functions hub card. Process Find, do not expose Three layers of internal visibility: work need-to-share, employment records need-to-know, directory with a mandatory core and a choice. One transparency policy for everything creates silos or glass employees. Process Visibility contract Operating artifact `visibility-contract.md`: three layers, work allowlist, record gate, mandatory vs optional core, preference record, revocation through the index, kill line. Practice — HR/privacy countersign. Privacy & protection Directory preference Steering of optional directory fields (photo, bio, skills, audience). Not consent and not a legal basis for people analytics or record access. Revocation must hit UI and index. Privacy & protection Employment record path The only allowed channel for contracts, pay, health, performance, and IDs: HRIS plus a named record site. Search, Copilot, and general marts do not inherit the record. Shadow copies count. Process Serial grain One row per physical device with a stable ID — not quote SKU quantity and not a CMDB dump. Process Install-base Shipped, still-active devices with location and contract — not the CRM product catalog and not quote quantity. Process Deal registration Protection claim in the vendor portal with a source ID. Registered is not a CRM pipeline stage and not a booking. Process Rebate vs booking A manufacturer or distributor payment (rebate/SPIFF) is not a booking, not hardware margin, and not maintenance ARR. Process RMA Return or replacement on the serial — not on the ticket alone and not on account 360. Process Field-service SLA Service clock on the serial (and site), not on the account. A ticket without a serial is not an SLA case. Architecture Medallion Layer model for raw, cleaned, and consumable data — a storage pattern, not a full governance architecture and not a substitute for owner, contract, and purpose. Process Codetermination DACH C-gate: the works council is Consulted on employee data, not the Data Owner. Evidence before processing; a return path and scope are mandatory. Process Trust claim Which consumer decision becomes more reliable in 90 days, and how the person notices it. Catalog coverage does not count. No trust claim, no Placement. Process Placement Default Field-Sales SKU: partial scope on a live initiative without owning the stack. A pilot needs a consumer; a platform lane needs a governance owner. Process SKU One deliverable service cell: Placement, Pilot, platform lane, or handover — one, not four. Sales does not design Ranger, dbt, or a semantic layer as a giveaway.

Buzzword Quiz

Gemischte Single- und Multi-Choice-Fragen aus dem Glossar — neue Begriffe fließen automatisch ein.

Kategorien

Optional — ohne Auswahl spielen alle Kategorien mit.

Fachbereiche

Optional — ohne Auswahl spielen alle Fachbereiche mit.

Tour