Salesforce Certified Data Architecture and Management Designer (SU18)
Difficulty: Intermediate | Prerequisites: Understanding of Salesforce objects and relationships (see the Data Modeling study notes), basic familiarity with data migration concepts.
Data quality, governance, and migration are tightly linked topics on the Data Architecture exam. Poor data quality (duplicates, incomplete records, incorrect field values) breaks dashboards and reporting. Governance covers the processes and tools for keeping data clean over time. Migration covers the mechanics of moving data between Salesforce orgs while preserving relationships and record integrity. Across all three areas, external IDs are the recurring thread.
Duplicate records
Two or more records in Salesforce that represent the same real-world entity (person, company, transaction). Duplicates cause inflated counts in reports, wasted sales effort, and inconsistent customer experiences.
In simple terms, if a customer has two Contact records, two different sales reps might call them without knowing about each other's call.
Data stewardship
The practice of managing and overseeing an organisation's data assets to ensure quality, consistency, and proper use. In a Salesforce context, data stewardship involves reviewing metadata for redundant fields, running reports to identify required fields, and reviewing the sharing model.
Validation rule
A Salesforce feature that enforces data entry standards by defining conditions that must be true before a record can be saved. If the condition is not met, the save is blocked and an error message is shown.
In simple terms, validation rules are the guardrails that stop users (and integrations) from saving incomplete or incorrect data.
External ID
A custom field marked as a unique identifier from an external system. External IDs are indexed for fast lookup, support upsert operations, and are critical for cross-system data matching and migration.
In simple terms, it is the field that says "this Salesforce record is the same entity as record 12345 in the ERP."
Master Data Management (MDM)
A discipline (and often a dedicated system) for maintaining a single, authoritative source of truth for key business entities like customers, products, and locations across all systems in an organisation.
In simple terms, the MDM system is the golden record. Every other system, including Salesforce, references its customer ID.
Data quality dashboard
A Salesforce dashboard designed to monitor data completeness, accuracy, and duplication. It surfaces records that are incomplete or incorrect so data stewards can take action.
Picklist field
A Salesforce field type that restricts input to a predefined set of values, preventing free-text entry errors. Using picklists over free text fields is a fundamental data quality technique.
Duplicate records are one of the most common data quality problems in Salesforce, and the exam tests your ability to identify them as the root cause of business problems.
Exam scenario
Customers are being called by multiple sales reps, and the second rep is unaware of the first rep's call. The data quality problem is duplicate Contact records. When the same customer exists as two separate Contact records, each owned by a different rep, neither rep sees the other's activity history.
Addressing data quality issues for dashboards
When executive dashboards are unreliable because of incomplete records, incorrect field entries, and duplicate entries, the exam expects three recommendations:
Use third-party data providers to enrich and augment information entered in Salesforce (fills gaps in incomplete records).
Use Salesforce features such as validation rules to prevent incomplete and incorrect records at the point of entry.
Design and implement data quality dashboards to monitor and act on records that are incomplete or incorrect.
Why the other options are less appropriate
Periodically exporting data for cleansing and re-importing is reactive and error-prone, not a sustainable solution.
Building a separate data warehouse for executive reporting does not fix the underlying data quality problem in Salesforce.
Before proposing design recommendations for a Salesforce org, a data architect should review three key areas.
Three areas to review
Review metadata XML files for redundant fields to consolidate. Over time, orgs accumulate fields that overlap in meaning or are no longer used. Identifying and consolidating these is a foundational stewardship activity.
Run key reports to determine what fields should be required. Reports reveal which fields are consistently left blank and which ones matter for business decisions. This informs validation rule design.
Review the sharing model to determine impact on duplicate records. The sharing model affects visibility. If duplicate records are owned by different users or sit in different sharing groups, the duplication problem may be invisible to individual users but compound in reports.
Why the other options are less appropriate
Determining whether integration points create records is useful context, but it is not one of the three primary review areas for a stewardship engagement.
Exporting the setup audit trail tells you what changed and when, but does not directly reveal which fields are in active use or which should be required.
When a public website creates Lead records in Salesforce via the REST API, maintaining data quality at the point of entry is critical.
Two recommended techniques
Client-side validation of phone number and email formats. Catching format errors before submission reduces junk data flowing into Salesforce.
Prefer picklist fields over free-text fields where possible. Picklists constrain input to known values, eliminating misspellings and inconsistent entries.
Why the other options are not the answer
HTTPS ensures data is encrypted in transit. That is a security concern, not a data quality technique.
Using cookies to track multiple form submissions addresses duplicate lead creation, which is a valid concern but is not framed as a "data quality" technique in the same way that validation and picklists are.
When a customer's data lives across multiple systems (CRM, billing, MDM, marketing, contract management), uniquely identifying the customer across all of them is a core data architecture challenge.
Recommended approach
Create a custom field as an external ID on the relevant Salesforce object (e.g. Account or Contact) to hold the customer ID from the MDM solution. The MDM system serves as the authoritative source for the customer's unique identity, and Salesforce references it.
Why external ID from the MDM, not the Salesforce ID
The Salesforce ID is specific to the Salesforce org. Storing it in every external system couples those systems to Salesforce's internal identifiers, which breaks if you migrate orgs or if a system does not integrate directly with Salesforce.
The MDM system's customer ID is designed to be the cross-system identifier. Referencing it in Salesforce via an external ID field keeps the architecture clean.
Why the other options are less appropriate
Creating a custom cross-reference object adds unnecessary complexity. The external ID field on the existing object does the same job more simply.
Creating a separate customer database duplicates the MDM system's role.
Real-world application
This pattern is standard in any enterprise where Salesforce is one of several systems. The call centre agent sees the MDM customer ID on the Account record, and every other system (billing, marketing, support) can use that same ID to look up the customer.
When migrating data from one Salesforce org to another, preserving the relationship hierarchy (master-detail and lookup relationships) is the central challenge. Record IDs are unique to each org, so you cannot simply export and import with the same IDs.
Three steps to maintain relationships during migration
Use Data Loader to export from the source org, then import or upsert into the target org in sequential order. Parent records must be loaded before child records so that the parent IDs exist when the children reference them.
Create an external ID field on each object in the target org and map the source record IDs to this field. This lets you use upsert to match records by their source org ID rather than the target org's new IDs.
Replace source record IDs with new record IDs from the target org in the import file. Once parent records are loaded and have new IDs in the target org, update the child import files to reference those new parent IDs.
Why the other options are less appropriate
Redefining master-detail relationships as lookups in the target org would change the data model's behaviour (losing cascade delete, sharing inheritance, and roll-up summaries). That is a structural change, not a migration step.
Keeping the source record IDs in the relationship fields of the import file will not work because those IDs do not exist in the target org.
Sequencing matters
The load order for a migration with parent-child relationships should follow the dependency chain:
Load parent objects first (e.g. Accounts).
Load child objects next (e.g. Contacts), referencing the parent's external ID.
Load grandchild objects (e.g. Cases), referencing the child's external ID.
If master-detail relationships exist, this order is mandatory because the child record cannot be saved without the parent.
Students often think duplicate Contacts are a minor annoyance. On the exam, they are the root cause of serious business problems like customers being called by multiple reps.
Students sometimes assume that building a data warehouse fixes data quality in Salesforce. It does not. The warehouse reports on the same bad data unless the source is cleaned.
Students confuse HTTPS (a security measure) with a data quality technique. HTTPS protects data in transit; it does nothing to validate the content of a form submission.
Students often try to keep source org record IDs in relationship fields during migration. Those IDs are meaningless in the target org. External IDs are the migration mechanism.
Duplicate management questions test whether you can connect a business symptom (reps calling the same customer) to a data quality root cause (duplicate Contacts).
Data stewardship questions list five plausible review areas and ask you to pick three. Remember: metadata review, key reports, and the sharing model.
Web form data quality questions pair two correct techniques (client-side validation and picklists) against two distractors (HTTPS and cookies).
The MDM/external ID question tests your understanding of cross-system identity. The MDM customer ID goes into a Salesforce external ID field.
Migration questions test sequencing (parents before children), external ID mapping, and the mechanics of replacing source IDs with target IDs.
True or false: Duplicate Activity records are the most likely cause of multiple reps calling the same customer. (False. Duplicate Contact records are the cause.)
Fill in the blank: To uniquely identify a customer across Salesforce and four other systems, create a custom field as an ______ to hold the customer ID from the MDM solution. (External ID.)
True or false: HTTPS is a data quality technique for web forms. (False. It is a security measure.)
Fill in the blank: During org-to-org migration, ______ records must be loaded before child records. (Parent.)
True or false: You can keep source org record IDs in lookup fields when migrating to a target org. (False. Source IDs do not exist in the target org.)
Q: Customers are being called by multiple sales reps who are unaware of each other's calls. What data quality problem causes this?
A: Duplicate Contact records exist in the system.
Q: Executive dashboards are unreliable due to incomplete records, incorrect field entries, and duplicates. Name three recommended steps.
A: Explore third-party data providers to enrich data; leverage validation rules to prevent incomplete and incorrect records; design and implement data quality dashboards to monitor and act on bad records.
Q: Which three areas should a data architect review before proposing stewardship recommendations?
A: Review metadata XML files for redundant fields, run key reports to determine what fields should be required, and review the sharing model to determine impact on duplicate records.
Q: Which two techniques help maintain data quality on web forms that create Lead records via the REST API?
A: Client-side validation of phone and email formats, and preferring picklist fields over free-text fields.
Q: A customer's data lives across five systems. How should a data architect uniquely identify the customer across all of them?
A: Create a custom field as an external ID in Salesforce to hold the customer ID from the MDM solution.
Q: What three steps maintain relationship hierarchies during org-to-org migration?
A: Use Data Loader to export and import/upsert in sequential order; create an external ID field on each object in the target org and map source record IDs to it; replace source record IDs with new target org IDs in the import file.
Data quality ties directly to reporting and analytics: bad data in, bad dashboards out. Validation rules connect to the broader Salesforce automation toolkit (workflow rules, Process Builder, flows). External IDs bridge this topic to the data migration section and to the MDM/cross-system identity pattern. The sharing model review in data stewardship connects to the security and access topics covered in the Salesforce Identity and Access Management Designer exam.
Duplicate records, duplicate Contact, data quality, data stewardship, data governance, validation rule, required field, picklist field, client-side validation, REST API lead creation, MDM, master data management, external ID, cross-system identity, customer 360, data migration, org-to-org migration, Data Loader, upsert, sequential load, parent-child load order, data quality dashboard, third-party data enrichment, data cleansing, sharing model, metadata review