Salesforce Certified Data Architecture and Management Designer (SU18)
Difficulty: Intermediate | Prerequisites: Basic understanding of Salesforce objects, SOQL, and the Salesforce data model.
Salesforce provides several mechanisms for handling large data volumes and improving query performance: skinny tables for faster reads, PK Chunking for exporting tens of millions of records, big objects for storing 100 million+ rows, and the Bulk API with GZIP compression for efficient data movement. Knowing when to reach for each tool, and why, is a core part of the Data Architecture and Management Designer exam.
Skinny table
A custom, read-only table maintained by Salesforce Support that contains a subset of fields from a standard or custom object. Skinny tables sit alongside the main database tables and are optimised for fast reads by stripping out soft-deleted records and avoiding joins.
In simple terms, think of it as a slimmed-down copy of your object table that Salesforce keeps in sync for you, so reports and queries on large objects run faster.
PK Chunking (Primary Key Chunking)
A Bulk API feature activated via the Sforce-Enable-PKChunking header. It splits a large query job into smaller chunks based on record ID ranges, preventing full-table-scan timeouts on very large objects.
In simple terms, instead of trying to read all 40 million records in one go, PK Chunking breaks the job into manageable bites keyed off the record ID.
Big object
A Salesforce object type designed to store and manage very large datasets (hundreds of millions of records or more). Big objects support only a subset of features compared to standard or custom objects and are created via the Metadata API or Salesforce DX, not through point-and-click Setup.
In simple terms, big objects are Salesforce's answer to "we need to keep a billion rows of historical data without blowing up our org's storage or performance."
Bulk API
A REST-based API optimised for loading or extracting large sets of data. It processes records asynchronously in batches and is the recommended tool for data loads and exports exceeding a few thousand records.
In simple terms, the Bulk API is what you use when Data Loader or the SOAP API would time out or crawl.
GZIP compression
A data-compression technique applied to API responses (and sometimes requests) to reduce payload size over the wire. In the context of Salesforce data exports, enabling GZIP compression helps avoid timeouts when extracting large record sets through the Bulk API.
In simple terms, it squeezes the data smaller so it travels faster between Salesforce and your external system.
Skinny tables are a performance optimisation feature that Salesforce Support creates on request. They copy a flat subset of an object's fields into a separate table that is faster to query.
Why they are fast
They exclude soft-deleted records, so the query engine never wastes time filtering out rows in the recycle bin.
They avoid resource-intensive joins. Standard Salesforce queries often join multiple underlying database tables (e.g. to pull in custom fields stored separately). A skinny table collapses those into one flat structure.
They are kept automatically in sync with the source object. When records in the original object are created, updated, or deleted, the corresponding skinny table rows are updated by the platform.
Limitations to remember
Skinny tables support a maximum of 100 columns.
They can only contain fields from the object itself, not fields from related (parent/child) objects.
You cannot create or manage them through Setup. You must contact Salesforce Support.
They do not include soft-deleted records, which means queries against a skinny table will not return records sitting in the recycle bin.
When to use them
Consider requesting a skinny table when you have a high-volume object (millions of records) and reporting or SOQL queries against it are timing out, particularly if those queries only need a subset of the object's fields.
When exporting very large datasets (tens of millions of records) from Salesforce, a standard Bulk API query can fail with a full-table-scan timeout. PK Chunking solves this.
How PK Chunking works
You add the Sforce-Enable-PKChunking header to your Bulk API job.
Salesforce automatically splits the query into multiple batches, each covering a range of record IDs.
Each batch runs independently and returns its own result set, avoiding the single-query timeout.
Exam scenario to know
A company is exporting 40 million Account records using an ETL tool (e.g. Informatica Cloud). The query log shows a full-table-scan timeout. The correct recommendation is to enable PK Chunking via the Sforce-Enable-PKChunking header on the export job.
Other options that appear as distractors on the exam:
"Export-in-Parallel" is not a real Salesforce header.
Adding standard index fields to the query can help with selective queries, but does not solve a full-table-scan timeout on a 40-million-row export.
Adding a LIMIT clause with a batch size of 10,000 would require manual pagination logic and is not the recommended approach for this volume.
Big objects are purpose-built for storing and accessing very large datasets, typically 100 million records or more, within Salesforce.
How to create big objects
There are two supported methods:
Metadata API - define the big object in XML metadata and deploy it.
Salesforce DX (SFDX) - define the big object in your project source and push it to the org.
You cannot create big objects through the standard Setup UI. The "Big Object" option in Setup is for viewing existing big objects, not for creating new ones. Likewise, Object Manager does not support creating big objects.
Key characteristics
Big objects do not support all the features of standard or custom objects (e.g. triggers, most standard UI, full SOQL).
They are queried using async SOQL.
They are designed for archival, audit, or historical tracking use cases where the data volume far exceeds what custom objects handle comfortably.
Real-world application
A company that needs to retain years of transaction history, IoT sensor readings, or audit logs at massive scale would use big objects rather than pushing hundreds of millions of rows into custom objects, which would degrade org performance.
When extracting large record sets for external systems (e.g. a Business Intelligence platform), choosing the right API and configuration matters.
Bulk API
The Bulk API is designed for asynchronous processing of large data jobs. It handles both imports and exports and processes records in batches.
GZIP compression
For exports of around 1 million records, enabling GZIP compression on the Bulk API response reduces the data volume transmitted and helps avoid timeouts during the export process.
Exam scenario to know
A company wants to extract 1,000,000 Contact records for an external BI system. The recommended approach to avoid timeouts is to use GZIP compression. The SOAP API would struggle at this volume, and scheduling a Batch Apex job is not the standard pattern for BI data extraction.
Choosing the right approach by volume
Thousands of records: SOAP API or REST API are fine.
Hundreds of thousands to low millions: Bulk API with GZIP compression.
Tens of millions: Bulk API with PK Chunking enabled.
Hundreds of millions (archival): consider whether big objects or external storage are more appropriate than live extraction.
Students often think skinny tables can include fields from related objects (e.g. parent Account fields on a Contact skinny table). They cannot. Skinny tables only hold fields from the single object they are built on.
Students often confuse PK Chunking with generic parallelism. "Export-in-Parallel" is not a valid Salesforce API header. PK Chunking is the specific mechanism that splits a Bulk API query by ID range.
Students sometimes assume big objects can be created through Setup or Object Manager. They cannot. You must use the Metadata API or Salesforce DX.
Students sometimes reach for the Bulk API when GZIP compression on the standard Bulk API call would solve the timeout. The two are complementary, not alternatives: GZIP is a setting on the Bulk API, and PK Chunking is a different setting on the same API.
The exam tests whether you can pick the right tool for the right data volume. Skinny tables, PK Chunking, big objects, Bulk API, and GZIP compression each solve a different scale of problem.
Expect scenario questions framed around "Universal Containers" or a similar company hitting timeouts or needing to store massive record counts. The correct answer turns on matching the volume and use case to the right feature.
Know the three reasons skinny tables are fast (no soft-deleted records, no joins, kept in sync). This is a frequently tested "choose three" question.
Know that big objects require Metadata API or Salesforce DX to create. Point-and-click options are distractors.
Know that PK Chunking is enabled via a header (Sforce-Enable-PKChunking), not by modifying the SOQL query itself.
True or false: Skinny tables can contain fields from parent objects. (False.)
True or false: PK Chunking is enabled by modifying the SOQL WHERE clause. (False. It is enabled via a job header.)
Fill in the blank: Big objects are created using the ______ or Salesforce DX. (Metadata API.)
True or false: GZIP compression is an alternative to the Bulk API. (False. GZIP compression is a setting applied to Bulk API responses.)
Fill in the blank: Skinny tables support a maximum of ______ columns. (100.)
Q: What three characteristics make skinny tables fast?
A: They do not include soft-deleted records, they avoid resource-intensive joins, and their tables are kept in sync with the source tables when the source tables are modified.
Q: An ETL tool is exporting 40 million Account records from Salesforce and fails with a full-table-scan timeout. What is the recommended solution?
A: Modify the export job header to specify Sforce-Enable-PKChunking.
Q: Which two tools should a data architect use to create a custom big object?
A: The Metadata API and Salesforce DX (SFDX). Big objects cannot be created through Setup or Object Manager.
Q: A company wants to extract 1,000,000 Contact records for a BI system. What should be recommended to avoid timeouts?
A: Use GZIP compression on the Bulk API export.
Skinny tables and PK Chunking both relate to the broader topic of Large Data Volume (LDV) management in Salesforce, which also includes indexing strategies, query selectivity, and data archiving. Big objects connect to the data lifecycle and archival strategy topics. The Bulk API appears again in data migration scenarios (covered in the Migration and Integration study notes).
Skinny table, PK Chunking, Sforce-Enable-PKChunking, primary key chunking, big object, custom big object, Metadata API, Salesforce DX, SFDX, Bulk API, GZIP compression, large data volume, LDV, full table scan, query timeout, data export, ETL, Informatica Cloud, async SOQL, data archival, record storage limits