Source: AP Computer Science Principles Course Framework
Tags: binary, bits, bytes, data compression, lossless, lossy, analog, digital, sampling, abstraction, metadata, data mining, correlation, causation, number base, hexadecimal, run length encoding
Difficulty: Beginner to Intermediate Prerequisites: Basic arithmetic (powers of 2, division with remainders).
Big Idea 2 is about how computers store, represent, and process data. It accounts for 22% of the exam. Everything a computer does comes down to binary (0s and 1s), so understanding number systems is foundational. From there, the unit covers how real-world (analog) data gets converted to digital form, how files are compressed to save space, and how large data sets are cleaned, analysed, and visualised. The distinction between correlation and causation is a favourite exam topic. If you can convert between binary and decimal, explain lossless vs. lossy compression, and identify what metadata is (and is not), you are well prepared.
Computers represent all data in binary. Analog data from the real world is converted to digital through sampling, which is a form of abstraction. Data compression reduces file sizes using lossless (no data lost) or lossy (some data permanently removed) methods. Large data sets require cleaning, can reveal patterns through data mining, and correlation between variables does not imply causation.
Data
A collection of facts. In computing, data is stored and processed as binary digits.
Number base
The number of digits or digit combinations a system uses to represent values. Decimal is base 10; binary is base 2.
Decimal system (base 10)
Uses digits 0 through 9. Each position represents a power of 10 (ones, tens, hundreds, thousands).
Binary system (base 2)
Uses only 0 and 1. Each position represents a power of 2 (1, 2, 4, 8, 16, 32, etc.).
Bit
The smallest unit of information stored or manipulated on a computer. A single 0 or 1. Think of it as a light switch: on or off, true or false.
Byte
A group of 8 bits. One byte can represent 256 unique values (2^8), from 0 to 255.
Analog data
Data that is measured continuously. It changes smoothly over time, like the volume of music or the position of a runner. There are infinite possible values between any two points.
Digital data
A discrete, simplified representation of information using a finite set of values. Leaves out some detail compared to analog. The timestamp on a YouTube video is digital; the continuous flow of a live event is analog.
Sampling
Recording an analog signal at regular discrete intervals and converting those measurements into digital signals that can be stored. The individual measurements are called samples.
Data abstraction
Filtering out specific details to focus on the information needed to process the data. Using digital data to approximate real-world analog data is a form of abstraction.
Data compression
A set of steps for packing data into a smaller space while allowing the original data (or an approximation of it) to be recovered.
Lossless data compression
Compression that reduces file size without sacrificing any original data. The original can be perfectly reconstructed from the compressed version. Run length encoding is an example. Used mainly for text and programs where every bit matters.
Lossy data compression
Compression that sacrifices some data to achieve greater file size reduction. The original can only be approximately reconstructed. Used mainly for images, audio, and video where small losses are acceptable. Examples include converting to greyscale or lowering resolution.
Run length encoding (RLE)
A lossless compression technique that replaces repeating data with a count and the repeated value. For example, FFFFFIIIIIIVVVVVVVEEEE becomes 5F6I7V4E.
Redundancy (in data)
Repeated information within data. The more redundancy, the more effective compression can be.
Metadata
Data about data. It does not affect the data itself. Changes to metadata do not change the primary data. Used to help find and organise information (e.g. file name, date created, author, file size).
Big data
Very large data sets that are difficult to process with a single computer.
Correlation
A statistical relationship between two or more variables, measured from -1 (perfect negative) to +1 (perfect positive). A correlation of 0 means no relationship.
Outlier
A data point that significantly deviates from the overall pattern or trend.
Data mining
The process of examining very large data sets to find useful information such as patterns and trends.
Cleaning data
The process of making a data set uniform and consistent by fixing formatting issues, removing duplicates, and handling missing or invalid values.
Decimal (base 10): positions are powers of 10 (1, 10, 100, 1000, ...)
Binary (base 2): positions are powers of 2 (1, 2, 4, 8, 16, 32, 64, 128, ...)
To convert binary to decimal: multiply each digit by its position's power of 2 and add them up
Example: binary 0101 = (0 x 8) + (1 x 4) + (0 x 2) + (1 x 1) = 5
Largest value representable with n bits: 2^n - 1
Example: 8 bits can represent values up to 2^8 - 1 = 255
Total number of unique values with n bits: 2^n
Example: 8 bits = 2^8 = 256 unique values (0 through 255)
8 bits = 1 byte = 256 combinations
Colour values like [255, 255, 255] use 3 bytes (24 bits) of data
Analog data is continuous and smooth (the position of a runner over time, the sound of a voice)
Digital data is discrete, with a finite number of possible values (a YouTube video's timestamp showing 10:00 rather than the exact millisecond)
Converting analog to digital uses sampling: measure the signal at regular intervals
This conversion is a form of abstraction because it leaves out detail between samples
More frequent sampling = better approximation, but more storage required
Compression matters because without it, a three-minute song would be over 100 MB
Compression depends on two things: the amount of redundancy in the data, and the compression method used
Lossless: original data is perfectly recoverable. Used for text, code, and anything where a single changed bit would break the file
Run length encoding is a lossless technique
Inverting pixel colours and brightness values is a reversible (lossless) operation
Lossy: some data is permanently removed. Used for images, audio, and video
Examples: converting to greyscale, lowering resolution, reducing audio quality
Removed data is gone forever
Higher compression = smaller file, but more quality loss
Fewer bits does not necessarily mean less information
The actual size reduction depends on both the redundancy in the original and the algorithm used
Programs such as spreadsheets help organise data and find trends
Data transformation examples: modifying every element (e.g. converting litres to millilitres), filtering by category (e.g. students in a specific extracurricular), combining or comparing data across sources
Data visualisation tools: bar charts (categorical data), scatter plots (relationship between two numeric variables, uses Pearson r), line graphs (values over time), histograms (frequency distribution across ranges)
Large data sets may require parallel computing systems to process
Bias in data is created by the type or source of data being collected, and is not eliminated by simply collecting more data
Largest value with n bits: 2^n - 1
Number of unique values with n bits: 2^n
Binary to decimal: sum of (each digit x its power of 2)
Example: What is the largest value you can represent with 7 bits? 2^7 - 1 = 127
MP3, MP4, and JPG files all use data compression to keep file sizes manageable for storage and transmission.
Streaming services use lossy compression for video and audio to reduce bandwidth, which is why a low-quality stream looks blurry (pixels and sound data have been removed).
Students often think correlation implies causation. Two variables can be strongly correlated without one causing the other. Ice cream sales and drowning rates both rise in summer, but ice cream does not cause drowning.
Students sometimes believe metadata affects the data itself. It does not. Changing a photo's metadata (file name, date tag) does not alter the image.
Students confuse the "largest value" and "number of values" calculations. With 8 bits, the largest value is 255, but the total number of values is 256 (because you count from 0).
Students think lossy compression is always worse than lossless. Lossy is appropriate when small quality reductions are acceptable and much smaller file sizes are needed (images, music, video).
⚠️ Know how to convert between binary and decimal in both directions.
⚠️ Be able to calculate the largest value and the total number of values for a given number of bits.
⚠️ Lossless vs. lossy compression: know the trade-offs, examples of each, and when each is appropriate.
⚠️ Correlation does not equal causation. This is tested directly and often.
⚠️ Metadata does not change the primary data. Expect a question that tries to trick you on this.
⚠️ Understand what data abstraction means and recognise examples (digital approximation of analog, using a list to represent a data set).
True or False: 8 bits can represent 255 unique values.
Fill in the blank: The process of recording an analog signal at regular intervals is called __________.
True or False: Lossy compression allows you to perfectly reconstruct the original data.
Fill in the blank: Data about data is called __________.
True or False: Correlation between two variables means one causes the other.
Answers: 1. False (8 bits can represent 256 unique values, from 0 to 255). 2. Sampling. 3. False (lossy compression only approximates the original; some data is permanently lost). 4. Metadata. 5. False (correlation does not imply causation).
Q: What is the largest decimal value that can be represented using 10 bits?
A: 2^10 - 1 = 1023.
Q: A music file is compressed so that frequencies above 20,000 Hz are removed, since most humans cannot hear them. Is this lossless or lossy compression? Explain.
A: This is lossy compression. Data (the high-frequency sounds) is permanently removed and cannot be recovered. The original audio can only be approximately reconstructed.
Q: A student notices that cities with more fire stations tend to have more fires. They conclude that fire stations cause fires. What is wrong with this reasoning?
A: The student is confusing correlation with causation. Larger cities have both more fire stations and more fires because they have more buildings and people. The underlying variable (city size) drives both, but fire stations do not cause fires.
Q: Convert the binary number 11001010 to decimal.
A: (1 x 128) + (1 x 64) + (0 x 32) + (0 x 16) + (1 x 8) + (0 x 4) + (1 x 2) + (0 x 1) = 128 + 64 + 8 + 2 = 202.
Q: What is the difference between data and metadata? Give an example.
A: Data is the actual content (e.g. the pixels in a photograph). Metadata is information about that data (e.g. the date the photo was taken, the file size, the camera model). Changing or deleting metadata does not alter the photograph itself.
Binary representation connects directly to Big Idea 3 (Algorithms and Programming), where you work with data types, variables, and Boolean values (true/false, which map to 1/0).
Data abstraction here is the same principle behind procedural abstraction in Big Idea 3, where functions hide complexity.
Bias in data sets connects to Big Idea 5 (Impact of Computing), where computing bias and the digital divide are examined at a societal level.
binary, decimal, base 2, base 10, bit, byte, analog data, digital data, sampling, data abstraction, data compression, lossless compression, lossy compression, run length encoding, RLE, metadata, big data, correlation, causation, outlier, data mining, cleaning data, data transformation, data visualisation, bar chart, scatter plot, line graph, histogram, AP CSP, AP Computer Science Principles, Big Idea 2