SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
Data Preparation for ML and AI

Data Quality Validation, Bias Metrics and Class Imbalance for ML (MLA-C02)

20 min readMLA-C02 · Data Preparation for ML and AIUpdated

Validating data quality and managing bias means proving that training data is complete, correct, consistent and representative before a model learns from it, using AWS Glue Data Quality and AWS Glue DataBrew for automated checks, labeling controls for trustworthy targets, bias metrics and careful splitting for fair representation, and cleaning for outliers, gaps and duplicates. This MLA-C02 lesson covers Task 1.3 of the AWS Certified Machine Learning Engineer - Associate exam: data quality rules, labeling, bias metrics across numeric, text and image data, class imbalance, prompt-response validation and content safety screening for foundation-model training data.

What you’ll learn
  • Choose between AWS Glue Data Quality on the Data Catalog, Glue Data Quality in ETL jobs and AWS Glue DataBrew rulesets, and write DQDL rules with thresholds, where clauses and dynamic rules.
  • Design labeling for measurable quality with multiple annotators, gold-standard items, audit samples and model-assisted pre-labeling.
  • Choose randomized, stratified, ordered or key-based splits, shuffle before positional splits and apply label-preserving augmentation to the training split only.
  • Interpret pre-training bias metrics (CI, DPL, KL, JS, TVD, KS, CDD) and apply them to numeric, text and image data by attaching facets.
  • Resolve class imbalance with resampling after the split or with class and instance weights when the data cannot change.
  • Validate prompt-response pairs for Amazon Bedrock fine-tuning and screen training data with Bedrock Guardrails, Amazon Comprehend and Amazon Rekognition.
  • Clean data by treating outliers, sentinel codes, missing values and fuzzy duplicates without leaking test information.

How do you validate data quality with AWS Glue Data Quality?

AWS Glue Data Quality validates data by evaluating a ruleset written in Data Quality Definition Language (DQDL) and reporting a pass or fail per rule plus an overall score. It has two entry points, and the choice depends on whether the data is already at rest or still moving through a pipeline.

Data Catalog data qualityETL data quality (Evaluate Data Quality transform)
What it checksTables already registered in the AWS Glue Data Catalog (data at rest)Data as an AWS Glue ETL job reads and transforms it (data in transit)
Typical userData stewards and governance teams, from the console with no codeData engineers who own the pipeline
Getting startedRule recommendation runs analyse a table and propose a starting rulesetYou write the ruleset in the job (visual editor or script)
ResultsDataset-level rule outcomes and a score, on demand or on a scheduleDataset-level outcomes plus row-level results, so the job can stop, quarantine or route failing records in the same run
Reading only new dataFilter the run with a partition predicate (for example the newest load date)Use job bookmarks so each run processes only unprocessed data

Pick the Data Catalog route when tables already exist, nobody wants to build or maintain a job, and you want AWS to propose rules. Pick the ETL route when bad rows must be caught before they land: the transform adds per-row outcome columns, so the job can write passing rows to the training prefix and failing rows to a quarantine prefix. A check that runs after the data is written (a Catalog run, a DataBrew profile job) can report a bad batch but cannot keep its rows out. Job bookmarks are an ETL-job feature; they are not a setting on a Data Catalog evaluation run.

Glue Data Quality validates data on the way into training. Watching the data an endpoint receives in production is a different job (SageMaker Model Monitor, Domain 4).

How do you write DQDL rules with thresholds, conditions and history?

A DQDL ruleset is a list such as Rules = [ ... ], and the main design decision per rule is whether it tolerates any failures. Boolean rules pass only when every row meets the condition; ratio rules return a fraction that you compare with a threshold.

NeedBoolean rule (zero tolerance)Ratio rule (tolerance)
No empty valuesIsComplete "col" fails on a single nullCompleteness "col" >= 0.99 allows up to 1% nulls
No repeated valuesIsUnique "col" fails on a single duplicateUniqueness "col" >= 0.99 lets some duplicates through

Match the rule to the stated tolerance: a key column that must never repeat needs IsUnique, while a label that may be partly empty because those rows are dropped later needs Completeness with a threshold. Other common rule types include ColumnValues (allowed values or ranges), RowCount, ColumnExists and custom SQL.

Conditional rules with where clauses

When a check applies only to some rows, add a where clause (AWS Glue 4.0 and later), which filters the rows the rule evaluates: IsComplete "shipping_address" where "fulfilment_type = 'delivery'" fails if any delivery order lacks an address and ignores pickup orders. Read the clause carefully: it must filter on the condition column and check the target column, not the reverse.

Dynamic rules that follow history

A fixed threshold raises false alarms on data that naturally rises and falls, such as seasonal row counts. Dynamic rules compare a metric with its own recent history using last(k). last(1) is a single value, but last(k) with k greater than 1 returns several values, so it must be reduced with an aggregation: avg, median, min, max, sum, std or index. The aggregation encodes the business rule. For example, RowCount > avg(last(5)) * 0.8 fails only when today's count drops more than 20% below the average of the previous five runs. Comparing with min fails only when a value falls below every recent run, comparing directly with avg (no margin) fails on every ordinary below-average day, max sets the strictest bar, and index(last(k), 0) compares with the most recent run only. Choose the aggregation that matches how much normal variation the business accepts.

How does AWS Glue DataBrew validate data quality?

AWS Glue DataBrew validates data with a ruleset (data quality rules such as maximum percentage of missing values, duplicate rows or value ranges) that is evaluated by a profile job with the ruleset attached. DataBrew is the visual, no-code option, so it suits analysts who prepare datasets without writing code.

  • Profile jobs validate; recipe jobs transform. A recipe job applies cleaning steps and never evaluates a ruleset, so scheduling only recipe jobs produces no validation results. To check data nightly, schedule a profile job with the ruleset attached.
  • Act on the validation event, not the job state. A profile job finishes as SUCCEEDED even when rules fail, because the job itself ran correctly. The outcome of the checks is published to Amazon EventBridge as a DataBrew Ruleset Validation Result event with a validationState of SUCCEEDED or FAILED.
  • Gate downstream work on that event. An EventBridge rule matching validationState FAILED can notify Amazon SNS; a rule matching SUCCEEDED can start a SageMaker pipeline execution, so training runs only on data that passed.

Choosing between the two services: DataBrew rulesets fit analyst-owned datasets and visual preparation; AWS Glue Data Quality fits cataloged data lakes and pipelines that must handle failing rows inside an ETL job. The transformation side of DataBrew (recipes and feature engineering) belongs to Task 1.2.

How do you label and annotate data with measurable quality?

Labeling produces the ground-truth targets a supervised model learns from, and label quality comes from three controls: more than one judgment per item, items with known answers, and audits of a random sample. Errors in labels become errors the model learns, so these controls matter as much as volume.

  • Multiple annotators with consolidation improve each label. Sending every item to several workers and consolidating their answers (for example by weighted voting) fixes the errors that come from one person guessing on ambiguous items. It costs more per item.
  • Gold-standard items measure each annotator. Mixing items whose correct label is already known into the task stream gives a per-person accuracy score, so you can retrain or remove annotators who fall below a target. Consolidation alone improves labels but does not tell you who is accurate.
  • Random audit samples measure the dataset. Reviewing a random sample of finished labels gives an accuracy estimate for the whole set.
  • Model-assisted pre-labeling cuts cost. A foundation model (for example through Amazon Bedrock batch inference, which is priced below on-demand calls) proposes a label for every item, and human reviewers check a random sample to measure accuracy and expose prompt problems to fix. Unreviewed machine labels give no measure of quality.

Amazon SageMaker Ground Truth (existing customers)

Amazon SageMaker Ground Truth is no longer open to new customers; existing customers keep using it, without new features. A new account therefore needs another route, such as FM pre-labeling with human review. For teams that already use it, three facts are tested:

WorkforceWho labelsFits when
PrivateYour own employees or contractors, signed in through Amazon Cognito or an OIDC identity providerData must stay inside the organization (regulated or sensitive data, specialist staff)
VendorA labeling company from AWS Marketplace, working under confidentiality termsNo spare internal staff, but workers must be vetted
Public (Amazon Mechanical Turk)Anonymous independent contractorsNon-sensitive data, high volume, lowest cost

Annotation consolidation combines the answers of several workers per object, set by the number of workers per data object. Automated data labeling trains a model on the human answers and labels the objects it is confident about, sending the rest to people. It supports only single-label image classification, semantic segmentation, bounding box (object detection) and single-label text classification, and only batch jobs, not streaming jobs. Named entity recognition, multi-label classification and video tasks need human labelers. Automated labeling still depends on a human workforce to learn from.

How do dataset splitting, shuffling and augmentation reduce bias?

A split decides what the model is tested on, so a split that lets related rows or future information cross between training and test produces a score that will not hold in production. Choose the split that matches how the model will meet new data. SageMaker Data Wrangler (part of SageMaker Canvas) offers each of these as a transform.

SplitWhat it doesUse when
RandomizedAssigns rows to splits at randomRows are independent of each other
StratifiedKeeps the same class mix in each splitA rare class must appear in every split
OrderedKeeps row order, holding out the latest rowsTime series and forecasting: train on the past, test on the future
Split by keyKeeps all rows sharing a key (patient, customer, source document) in one splitSeveral rows describe the same entity

The symptom of a wrong split is a high test score that collapses on genuinely new data. Several voice recordings per speaker spread across both splits let a model recognise speakers rather than the words: split by key on the speaker ID. A shuffled split of hourly energy-demand readings lets the model train on hours after the test hours: use an ordered split. Stratifying or changing the ratio does not fix either leak.

Features can leak too. A trailing average that includes the value being predicted (for example a demand average whose window reaches the forecast hour) puts the answer inside the features. Compute rolling and lag features only from data available before the prediction point.

Shuffle before a positional split. If an export is sorted (for example by label), taking the first 80% for training and the last 20% for testing assigns splits by that column. Shuffle the rows first, or use a randomized split.

Augmentation for under-represented conditions

When a condition is rare in the data (low-light shots, motion blur, a regional accent), augmentation creates more training examples of it without new collection: brightness, contrast, noise, blur or synthetic weather for images; paraphrase or back-translation for text. Two rules keep it honest:

  • Transforms must preserve the label. Mirroring an image of handwritten letters turns b into d; rotating a digit can turn 6 into 9. Use only transforms that leave the meaning unchanged.
  • Augment the training split only. Measure improvement on held-out real examples of the target condition. A test set with synthetic effects measures how well the model handles the effects, not the real world. Duplicating the few real examples adds no variety.

Which pre-training bias metrics measure which kind of imbalance?

Pre-training bias metrics quantify imbalance in a dataset's facets (groups defined by a sensitive or relevant attribute, such as age group or gender) and labels before any model exists. They compare a favoured facet a with a disfavoured facet d. Amazon SageMaker Clarify computes them, but Clarify is no longer open to new customers (existing customers keep it), and every metric below is a simple formula you can reproduce with pandas or Spark.

MetricQuestion it answersReading the value
Class imbalance (CI)Is one facet under-represented in row count?(n(a) − n(d)) / (n(a) + n(d)), from −1 to 1; near 0 is balanced, a large positive value means facet d is rare
Difference in proportions of labels (DPL)Do the facets receive the positive label at different rates?q(a) − q(d), where q is the share of the facet with a positive label; near 0 means similar positive rates
KL divergence, Jensen-Shannon (JS)How different are the facets' label distributions?0 means identical; JS is symmetric and bounded
Lp-norm, total variation distance (TVD)Distance between label distributionsTVD is half the L1 distance
Kolmogorov-Smirnov (KS)Largest gap between the facets' label distributionsMaximum divergence across label values
Conditional demographic disparity (CDD)Does a gap remain after conditioning on a subgroup?Demographic disparity computed per subgroup and averaged by size

Keep the two headline metrics apart: CI is about how many rows each facet has, DPL is about how often each facet gets the positive outcome. A dataset can have a high CI and a DPL near zero (a representation problem but no outcome gap) or the reverse.

CDD addresses Simpson's paradox. If one group mostly applied for loan products that are rarely approved, its overall approval rate is low even when each product treats both groups alike; plain demographic disparity (DD) repeats that overall gap, while CDD with the loan product as the subgroup shows the within-product picture.

Post-training metrics such as difference in positive proportions in predicted labels (DPPL) need a trained model's predictions, so they cannot be used before training.

How do you apply bias metrics to numeric, text and image data?

Bias metrics work on a table of facet values and labels, whatever the underlying asset, so applying them to any data type means first attaching a facet attribute to every record.

  • Numeric facets: a continuous attribute such as income or years of tenure has one group per value until you define groups. Set a threshold (or bins) that matches the policy question, for example customers with five or more years of tenure against newer ones, then compute CI and DPL. Scaling the column does not create groups.
  • Text: the facet is often not stored. Derive it per document with a service or model, for example the dominant language of each support ticket with Amazon Comprehend, or the customer region from ticket metadata, then compute CI (is a group under-represented?) and DPL (does it carry a different share of positive labels?). Deriving a facet from the label itself, such as sentiment groups for a sentiment model, measures nothing.
  • Images: take the facet from annotations or metadata (for example annotated lighting condition, device model, capture site) and join it to each image's label. Pixel statistics such as brightness are a weak proxy when real annotations exist, and explaining a trained model (SHAP, saliency) answers a post-training question, not a data question.

Optimizing the distribution means acting on what the metrics show: collect more data for under-represented facets, augment them in the training split, resample or reweight (next section), and recompute the metrics to confirm the gap closed. The goal is a training set whose facet and label mix supports the model's use, evaluated per facet on real held-out data.

How do you resolve class imbalance without leaking test data?

Class imbalance (one label far rarer than the others) makes a model predict the common class and still score high accuracy; you resolve it either by changing the training data (resampling) or by changing how much each example counts in the loss (weighting).

TechniqueWhat it doesWatch out for
Random oversamplingDuplicates minority rowsExact copies can cause overfitting
Random undersamplingRemoves majority rowsThrows away data
SMOTECreates new minority rows by interpolating between nearby minority rowsNeeds numeric features to find neighbours
Class or instance weightsMakes rare-class errors cost more in the lossNo rows change, so it fits rules that forbid altering data

Split first, then resample the training split only. Oversampling or SMOTE before splitting puts copies (or interpolations) of the same rows in training and test, and makes the test class mix differ from production, so test recall looks excellent and production recall does not. The test set should keep the real class distribution.

Weighting when the data must stay as recorded

  • Binary XGBoost: scale_pos_weight multiplies the weight of positive examples; a typical value is sum(negative) / sum(positive), for example 285,000 / 15,000 = 19. A value below 1 does the opposite.
  • Multi-class XGBoost: scale_pos_weight is for binary problems. With the SageMaker AI built-in XGBoost and CSV input, add an instance-weight column directly after the label and set csv_weights to 1; without the flag the column is read as an ordinary feature.
  • Deep learning (text, image): pass class weights inversely proportional to class frequency to the cross-entropy loss (focal loss is another option). Weights proportional to frequency make the problem worse; label smoothing and learning-rate changes do not rebalance classes.

For text and image datasets, augmentation of the rare class (in the training split) is the data-side equivalent of SMOTE. Judge the result with per-class recall, precision or F1 on real data rather than overall accuracy.

How do you validate prompt-response pairs and screen training data for safety?

Validating AI training data means checking that every prompt-response pair has the right format and teaches the behaviour you want, and that no unsafe content enters the dataset. A fine-tuned model learns to produce the responses it sees, so a bad pair is a bad lesson.

Format and integrity checks for Amazon Bedrock fine-tuning

  • Schema: datasets are JSONL, one record per line. Text models use prompt and completion fields. Conversational models use the bedrock-conversation-2024 schema: a top-level system list, then messages that alternate user and assistant, beginning with a user turn. The system prompt does not go inside messages, and the response to learn belongs in the final assistant turn, not folded into the user turn.
  • Limits: records must fit the chosen model's token and record limits, which vary by model; check the model's quotas and drop or trim records that exceed them.
  • Content of the responses: remove pairs whose response is empty, boilerplate or a non-answer (such as "I can't help with that" or a reply that only pastes a link), because the model would learn to give non-answers.
  • Duplicates: collapse hundreds of near-identical pairs to a few representative ones so one behaviour does not dominate. Shuffling or lowering the learning rate leaves the flawed records in place.
  • Train/validation overlap: if pairs were paraphrased or derived from shared sources, split by source ID so related pairs stay in one file, and remove validation prompts that match training prompts. A validation loss near zero with poor answers to new questions is the symptom.

Content safety screening

ToolScreensFits when
Amazon Bedrock Guardrails ApplyGuardrail APIText, against a guardrail's content filters, denied topics, word filters and sensitive-information policiesYou need exactly the same policies as an existing guardrail, without invoking (and paying for) a foundation model
Amazon Comprehend DetectToxicContentText, returning an overall Toxicity score plus per-category scores (for example HATE_SPEECH, PROFANITY, INSULT, VIOLENCE_OR_THREAT)Generic toxicity screening; filter on the category the policy cares about
Amazon Rekognition DetectModerationLabelsImages (JPEG or PNG), returning moderation categories such as explicit or violent content; stored video uses the asynchronous StartContentModeration / GetContentModeration APIs insteadUnsafe material may be in the pixels themselves

Match the tool to the policy and the modality. If policy tolerates insults but forbids threats, filter on the VIOLENCE_OR_THREAT category score rather than tuning the overall Toxicity threshold, which cannot separate categories. A caption check does not see what is in an image, and general object detection (DetectLabels) is not a moderation check. Detecting and redacting PII is a separate step covered in Task 1.2.

How do you clean data: outliers, missing values and duplicates?

Cleaning data means treating outliers, missing values and duplicates so that each one is handled by what it actually is: an error, a placeholder or a genuine rare event. The same number can be any of the three, so read the scenario for what the values mean.

Outliers

Data Wrangler transformHow it sets the cut-offFits when
Min-max numeric outliersFixed lower and upper thresholds you supplyA known valid range exists (for example a physically possible engine temperature range)
Standard deviation numeric outliersMean ± k standard deviationsRoughly normal data with mild outliers
Robust standard deviation numeric outliersStatistics computed only from data between two quantilesExtreme values are large enough to inflate the mean and standard deviation and hide themselves
Quantile numeric outliersValues beyond chosen percentilesSkewed data with no known range

Each transform offers three fix methods: Remove the flagged rows, Clip values to the threshold, or Invalidate them, which marks them as missing so a later imputation step fills them. Do none of these to genuine extremes that carry the signal, such as traffic spikes that precede an outage. A sentinel code (a placeholder such as 9999 written when a sensor drops out) is not an outlier but a missing value: convert it to missing first, then impute. Data Wrangler also lists Replace rare under its outlier transforms, but it is the categorical option that groups infrequent categories and does not detect numeric outliers.

Missing values

  • Numeric, skewed: impute the median, which extreme values do not distort; the mean does.
  • Categorical: impute the mode (most frequent value). Averaging encoded categories is meaningless.
  • Missingness that predicts the target: add an indicator column recording that the value was missing, then impute, so the signal survives.
  • Drop rows only when few rows are affected and the gaps carry no signal; forward fill is for ordered data such as time series.
  • Fit on training data only. Compute the imputation statistic from the training split after splitting, store it, and apply the same value to the test set and to live inference requests, so evaluation is honest and serving matches training. Exclude sentinel codes when computing it.

Duplicates

Exact duplicates are removed by comparing rows. Fuzzy duplicates (the same customer spelled differently, with no shared ID) need entity matching: the AWS Glue FindMatches transform learns from labeled examples which records refer to the same entity. Tune it with the precision-recall setting: favour precision when false merges are costly (two patients merged into one), favour recall when missed duplicates are worse, and teach it with labeled matches and non-matches around the boundary. The accuracy-cost setting trades matching quality for compute cost. Deduplicate before splitting, or copies of one entity end up in both training and test.

Tip. Task 1.3 questions are scenarios with one deciding constraint in the middle: who owns the data and whether they write code, whether bad rows may reach training, whether rows may be changed, where the data must stay, whether the account is new, or what range the values can physically take. Options are often close variants of one another that differ in a single detail, so read each rule, setting or event name in full rather than recognising a familiar service. Several questions describe a symptom rather than naming the problem (a test score far above production performance, a suspiciously low validation loss, a metric that misses its own outliers) and ask for the fix, sometimes as a Select TWO where both halves of a solution are needed. Bias questions give metric values or a fairness question and ask which metric or conclusion applies; know what each pre-training metric measures and which metrics need a trained model's predictions.

Key takeaways
  • Glue Data Quality in an ETL job gives row-level results so bad rows can be quarantined in the same run; Data Catalog runs check data at rest, recommend rules and filter with partition predicates.
  • DQDL Boolean rules (IsComplete, IsUnique) allow zero failures; ratio rules (Completeness, Uniqueness) take a threshold; last(k) with k > 1 needs an aggregation such as min or avg.
  • A DataBrew ruleset is evaluated by a profile job, and the pass or fail lives in the Ruleset Validation Result event, not the job state.
  • SageMaker Ground Truth and SageMaker Clarify are closed to new customers; a new account labels with FM pre-labeling plus human audit and computes bias metrics itself.
  • Split by key for repeated entities, use an ordered split for time series, shuffle sorted exports and deduplicate before any split.
  • CI measures how many rows each facet has; DPL measures how often each facet gets the positive label; CDD conditions on a subgroup to expose Simpson's paradox.
  • Resample only the training split after splitting; when rows must stay as recorded, use scale_pos_weight (binary), XGBoost instance weights with csv_weights=1 (multi-class) or inverse-frequency loss weights.
  • Augment rare conditions with label-preserving transforms in training only, and measure on real held-out examples.
  • ApplyGuardrail screens text against an existing guardrail without invoking a model; Comprehend gives per-category toxicity scores; Rekognition DetectModerationLabels screens images.
  • Impute skewed numbers with the median and categories with the mode, fitted on the training split; treat sentinel codes as missing and keep informative extremes.

Frequently asked questions

What is the difference between AWS Glue Data Quality and AWS Glue DataBrew data quality rules?

AWS Glue Data Quality evaluates DQDL rulesets either on tables in the AWS Glue Data Catalog (data at rest, with rule recommendations and schedules) or inside an AWS Glue ETL job, where row-level results let the job quarantine failing records in the same run. AWS Glue DataBrew is the visual, no-code tool: a DataBrew ruleset is evaluated by a profile job, and the pass or fail is published as a Ruleset Validation Result event in Amazon EventBridge. Use Glue Data Quality for data lakes and pipelines, DataBrew for analyst-prepared datasets.

Why does a DataBrew profile job succeed when its data quality rules fail?

A DataBrew profile job reports whether the job ran, not whether the data passed. When an attached ruleset fails, the job still finishes as SUCCEEDED, and the rule outcome is published separately as a DataBrew Ruleset Validation Result event with a validationState of SUCCEEDED or FAILED. To gate training or send alerts on data quality, match that event in Amazon EventBridge instead of the job state change.

What is the difference between class imbalance (CI) and DPL?

Class imbalance (CI) measures whether one facet, such as an age group, has far fewer rows than another; values near 0 mean balanced representation. Difference in proportions of labels (DPL) measures whether the facets receive the positive label at different rates. A dataset can be under-represented for a group (high CI) while that group gets positive labels at the same rate (DPL near 0), so the two metrics answer different questions and are often reported together.

Should you oversample or use SMOTE before or after splitting the data?

After. Split the data into training and test sets first, then oversample, undersample or apply SMOTE to the training split only. Resampling before the split places copies or interpolations of the same rows in both sets and changes the test class mix, so test scores look far better than production performance. The test set should keep the real class distribution.

Is SageMaker Ground Truth still available?

Amazon SageMaker Ground Truth is no longer open to new customers; existing customers can keep using it, but no new features are planned. Teams already on Ground Truth can still choose private, vendor or Mechanical Turk workforces, annotation consolidation and automated data labeling. A new account needs another approach, such as pre-labeling with a foundation model through Amazon Bedrock batch inference and auditing a random sample with human reviewers.

How do you screen fine-tuning data for unsafe content on AWS?

For text that must meet the same policies as an existing Amazon Bedrock guardrail, call the ApplyGuardrail API, which evaluates content filters and denied topics without invoking a foundation model. For generic toxicity, Amazon Comprehend DetectToxicContent returns an overall score and per-category scores such as HATE_SPEECH and PROFANITY, so you can filter on the categories your policy forbids. For images, Amazon Rekognition DetectModerationLabels returns moderation categories such as explicit or violent content.

Should missing values be imputed with the mean or the median?

Use the median for skewed numeric columns, because a few extreme values pull the mean but not the median, and use the mode for categorical columns. If the fact that a value is missing predicts the target, add a missing-indicator column before imputing. Compute the statistic on the training split only, store it, and apply the same value to the test set and to live inference requests.

How do you deduplicate records that have no shared ID?

Use fuzzy matching. The AWS Glue FindMatches transform learns from labeled examples which records refer to the same entity even when names and addresses are spelled differently. Tune its precision-recall setting toward precision when wrongly merging two different entities is the costly error, and toward recall when missed duplicates matter more. Deduplicate before splitting so one entity does not appear in both training and test data.

Source

This lesson covers the "Data Preparation for ML and AI" domain of the official MLA-C02 exam guide. Vendors revise their guides — check the source for the current version.

Test yourself on this topic
Practice questions with full explanations.
Practice now

Sign up free to mark lessons complete, bookmark topics and track your exam readiness.

Spotted a mistake in this lesson?