Three Types of Learning, Reinforcement Schedules, and Operant Conditioning, PSYCH 1100 Ch. 8 – Study Notes
offline

Source: The Adaptive Mind, pp. 279–280 (8-2), pp. 298–301 (8-4b), pp. 305–306 (8-4e)

Tags: classical conditioning, operant conditioning, observational learning, reinforcement, punishment, positive reinforcement, negative reinforcement, positive punishment, negative punishment, schedules of reinforcement, fixed ratio, variable ratio, fixed interval, variable interval, shaping, extinction, Skinner, Pavlov, Bandura, Psychology 1100, Ohio State

Difficulty: Intermediate Prerequisites: None strictly required, but familiarity with basic neurotransmitter concepts (especially dopamine and reward) from the Biological Mind chapter is helpful.


Big Picture

The learning chapter is one of the most heavily tested areas in introductory psychology, and three subsections appear on the "most missed" list: the overview of the three main types of learning, schedules of reinforcement, and applying operant conditioning. Students tend to confuse positive and negative reinforcement, mix up punishment and negative reinforcement, and struggle with which reinforcement schedule produces which pattern of behaviour. If you can draw the distinction between "positive/negative" (adding/removing) and "reinforcement/punishment" (increasing/decreasing behaviour), the rest falls into place.


TL;DR

Psychology recognises three main types of learning: classical conditioning (associating two stimuli), operant conditioning (learning from consequences), and observational learning (learning by watching others). Within operant conditioning, reinforcement schedules determine how consistently and quickly behaviours are learned and how resistant they are to extinction. Applying operant conditioning involves techniques like shaping, token economies, and behaviour modification.


Key Terms

Learning

A relatively permanent change in behaviour or knowledge resulting from experience. In simple terms, if your behaviour changes because of something that happened, you learned something.

Classical conditioning

Learning through the association of two stimuli, so that a previously neutral stimulus comes to trigger a response it did not trigger before. Pavlov's dogs learned to salivate at the sound of a bell because the bell was repeatedly paired with food.

Operant conditioning

Learning in which behaviour is strengthened or weakened by its consequences. If a behaviour is followed by something pleasant, it is more likely to be repeated; if followed by something unpleasant, it is less likely.

Observational learning (social learning)

Learning by watching the behaviour of others and noting the consequences. Bandura's Bobo doll experiment showed that children who watched an adult behave aggressively toward a doll were more likely to imitate that aggression.

Reinforcement

Any consequence that increases the likelihood of a behaviour being repeated. This is the defining feature: if the behaviour goes up, it was reinforced.

Positive reinforcement

Adding something desirable after a behaviour to increase its frequency. "Positive" means adding, not "good." A treat given to a dog after it sits is positive reinforcement.

Negative reinforcement

Removing something undesirable after a behaviour to increase its frequency. "Negative" means removing, not "bad." Taking aspirin removes a headache, which reinforces the behaviour of taking aspirin. The seatbelt buzzer stops when you buckle up, which reinforces buckling up.

Punishment

Any consequence that decreases the likelihood of a behaviour being repeated. If the behaviour goes down, it was punished.

Positive punishment

Adding something undesirable after a behaviour to decrease its frequency. A speeding ticket (adding a fine) after speeding is positive punishment.

Negative punishment

Removing something desirable after a behaviour to decrease its frequency. Taking away a teenager's phone (removing a privilege) after breaking curfew is negative punishment.

Shaping

Reinforcing successive approximations toward a target behaviour. Used when the desired behaviour is too complex to occur spontaneously. You reward each step closer to the goal until the full behaviour is achieved.

Extinction (operant)

The gradual weakening and disappearance of a learned behaviour when reinforcement is no longer provided. If a rat that has been pressing a lever for food stops receiving food, lever-pressing will eventually stop.

Schedule of reinforcement

The rule governing how often and under what conditions a behaviour is reinforced. Different schedules produce different patterns of responding and different resistance to extinction.

Continuous reinforcement

Reinforcing every instance of the target behaviour. Produces rapid learning but also rapid extinction when reinforcement stops.

Partial (intermittent) reinforcement

Reinforcing the target behaviour only some of the time. Produces slower learning but much greater resistance to extinction.

Fixed-ratio (FR) schedule

Reinforcement is delivered after a fixed number of responses. Example: a factory worker paid for every 10 items produced. Produces a high, steady rate of responding with a brief pause after each reinforcement (the "post-reinforcement pause").

Variable-ratio (VR) schedule

Reinforcement is delivered after an unpredictable number of responses. Example: slot machines pay out after a variable number of plays. Produces the highest and most consistent rate of responding, with no predictable pause, and is the most resistant to extinction.

Fixed-interval (FI) schedule

Reinforcement is delivered for the first response after a fixed period of time has passed. Example: checking the post once a day because delivery happens at the same time. Produces a "scalloped" pattern: responding is slow right after reinforcement and accelerates as the next reinforcement approaches.

Variable-interval (VI) schedule

Reinforcement is delivered for the first response after a variable period of time has passed. Example: checking your phone for messages, since notifications arrive at unpredictable times. Produces a slow, steady rate of responding.

Token economy

A system in which tokens (points, stickers, chips) are given as secondary reinforcers for desired behaviours and can later be exchanged for primary reinforcers (food, privileges, items). Used in classrooms, psychiatric facilities, and rehabilitation settings.

Partial reinforcement extinction effect (PREE)

The finding that behaviours reinforced on a partial schedule are harder to extinguish than those reinforced continuously. This is because the organism has already learned that reinforcement does not follow every response, so the absence of reinforcement is not immediately interpreted as "it is over."


Core Content

The Three Main Types of Learning

  • Classical conditioning (Pavlov):

    • Involves learning an association between two stimuli.

    • Before conditioning: food (unconditioned stimulus, US) produces salivation (unconditioned response, UR). A bell (neutral stimulus) produces no salivation.

    • During conditioning: bell is paired with food repeatedly.

    • After conditioning: bell alone (conditioned stimulus, CS) produces salivation (conditioned response, CR).

    • Key concepts: acquisition, extinction, spontaneous recovery, generalisation, discrimination.

  • Operant conditioning (Skinner):

    • Involves learning from the consequences of one's own behaviour.

    • The organism's behaviour is instrumental in producing an outcome (which is why it is sometimes called instrumental conditioning).

    • Core principle: behaviours followed by reinforcement increase; behaviours followed by punishment decrease.

  • Observational learning (Bandura):

    • Learning occurs by observing a model and noting the consequences of their behaviour.

    • Does not require direct reinforcement of the observer.

    • Bandura's Bobo doll experiment: children who watched an adult model punch and kick an inflatable doll later imitated the aggressive behaviour, especially if the model was rewarded or faced no consequences.

    • Four processes required: attention (notice the model), retention (remember the behaviour), reproduction (be able to perform it), motivation (have a reason to do it).

Schedules of Reinforcement

  • After a behaviour is established, the schedule on which it is reinforced determines how the organism responds and how resistant the behaviour is to extinction.

  • Two dimensions define the four main schedules:

    • Ratio vs. interval: is reinforcement based on the number of responses (ratio) or the passage of time (interval)?

    • Fixed vs. variable: is the requirement predictable (fixed) or unpredictable (variable)?

  • Fixed ratio (FR): reinforcement after every Nth response.

    • Pattern: high rate of responding, brief pause after reinforcement.

    • Example: piece-rate pay (paid per unit produced).

  • Variable ratio (VR): reinforcement after an unpredictable number of responses.

    • Pattern: highest and most consistent rate. No pauses.

    • Example: gambling (slot machines, lottery tickets). This is why gambling is so persistent.

    • Most resistant to extinction of all four schedules.

  • Fixed interval (FI): reinforcement for the first response after a set time period.

    • Pattern: "scalloped" response curve. Low responding right after reinforcement, increasing as the next interval approaches.

    • Example: studying increases as an exam approaches, then drops right after.

  • Variable interval (VI): reinforcement for the first response after an unpredictable time period.

    • Pattern: slow, steady rate of responding. No scallop.

    • Example: pop quizzes (you never know when the next one will come, so you maintain a steady level of preparation).

  • Key comparison: ratio schedules generally produce higher response rates than interval schedules, because faster responding directly leads to more reinforcement. Variable schedules produce more consistent responding and greater extinction resistance than fixed schedules.

Applying Operant Conditioning

  • Shaping: used to teach complex or novel behaviours that would not occur on their own.

    • The trainer reinforces successive approximations, each step closer to the final target.

    • Example: teaching a pigeon to turn in a circle. First reinforce any turning of the head, then a quarter-turn, then a half-turn, and so on.

    • Shaping is how animal trainers teach tricks and how therapists teach new skills to individuals with developmental disabilities.

  • Token economies: a structured reinforcement system.

    • Tokens serve as secondary reinforcers (they have no inherent value but acquire value through association with primary reinforcers).

    • Effective in managing behaviour in institutional settings, classrooms, and rehabilitation programmes.

  • Behaviour modification: the systematic application of operant conditioning principles to change behaviour.

    • Used in therapy (applied behaviour analysis for autism), education (classroom management), and self-improvement (habit tracking).

    • Steps typically include: identifying the target behaviour, establishing a baseline, selecting appropriate reinforcers, implementing the programme, and monitoring results.

  • Limitations of punishment:

    • Punishment suppresses behaviour but does not teach an alternative.

    • It can produce fear, avoidance, and aggression.

    • Effects are often temporary unless the punishment is consistently applied.

    • For these reasons, reinforcement of desired behaviour is generally preferred over punishment of undesired behaviour.


Real-World Applications

Variable-ratio schedules explain why gambling, social media scrolling, and checking email are so compelling: the unpredictable reward keeps you responding. Token economies are used in schools (classroom reward charts) and in cognitive rehabilitation. Shaping is the basis of animal training (from service dogs to marine mammal shows) and is also applied in therapeutic contexts for individuals learning new skills. Understanding the difference between reinforcement and punishment has practical implications for parenting, teaching, and management.


Common Misconceptions

  • "Negative reinforcement is the same as punishment." This is the most common error in the entire learning unit. Negative reinforcement increases behaviour (by removing something aversive). Punishment decreases behaviour. "Negative" refers to removal, not to something bad happening.

  • "Positive means good and negative means bad." In operant conditioning, "positive" means adding a stimulus and "negative" means removing one. A positive punishment adds something unpleasant; a negative reinforcement removes something unpleasant. The terms describe the operation, not a value judgement.

  • "Continuous reinforcement is the best schedule for maintaining behaviour." Continuous reinforcement produces the fastest learning but also the fastest extinction. Partial reinforcement is more resistant to extinction, making it better for maintaining behaviour long-term.

  • "Punishment is the most effective way to change behaviour." Punishment only suppresses behaviour temporarily and does not teach an alternative. Reinforcement of a desired replacement behaviour is generally more effective and has fewer side effects.


Why It Matters / Exam Flags

⚠️ The positive/negative × reinforcement/punishment grid is the single most important thing to have memorised for this section. Practice classifying scenarios into one of the four quadrants.

⚠️ Know the four schedules of reinforcement by name, definition, response pattern, and a real-world example for each.

⚠️ Variable-ratio produces the highest response rate and greatest extinction resistance. This comes up frequently.

⚠️ Be able to distinguish classical conditioning from operant conditioning from observational learning when given a scenario. The key question is: is the organism learning an association between stimuli (classical), learning from consequences of its own behaviour (operant), or learning from watching someone else (observational)?

⚠️ Know what shaping is and when it is used. A question might describe a training scenario and ask what operant technique is being applied.


Quick Self-Test

True or False: Negative reinforcement decreases the frequency of a behaviour.

A: False. All reinforcement (positive and negative) increases behaviour. Negative reinforcement increases behaviour by removing an aversive stimulus.

Fill in the blank: A ________ schedule of reinforcement delivers reinforcement after an unpredictable number of responses and produces the highest rate of responding.

A: Variable-ratio (VR).

True or False: In observational learning, the observer must be directly reinforced for learning to occur.

A: False. Observational learning can occur without direct reinforcement of the observer. The observer watches the model and notes the consequences to the model (vicarious reinforcement or punishment).

Fill in the blank: Reinforcing successive approximations toward a target behaviour is called ________.

A: Shaping.


Practice Q&A

Q: A child throws a tantrum in a shop and the parent buys them a toy to stop the screaming. Identify the operant conditioning principles at work for both the child and the parent.

A: The child is positively reinforced (receiving a toy increases the likelihood of future tantrums). The parent is negatively reinforced (the removal of the screaming increases the likelihood of buying a toy in future tantrums).

Q: A teacher gives pop quizzes at unpredictable times throughout the semester. What schedule of reinforcement does this represent, and what pattern of studying does it produce?

A: Variable-interval schedule. Students cannot predict when the next quiz will occur, so they maintain a slow, steady rate of studying rather than cramming right before a known test date.

Q: Explain why slot machines are so addictive using reinforcement schedule theory.

A: Slot machines operate on a variable-ratio schedule. The payout occurs after an unpredictable number of plays, so the player cannot know when the next win will come. This produces a high, consistent rate of responding (continued play) and extreme resistance to extinction (players continue even during long losing streaks).

Q: A dog trainer wants to teach a dog to roll over, but the dog has never done this before. What operant technique should the trainer use, and how?

A: Shaping. The trainer should reinforce successive approximations: first reward lying down, then reward lying down and turning slightly, then a half-roll, and finally a full roll-over. Each step closer to the target behaviour is reinforced until the complete behaviour is achieved.

Q: How is negative punishment different from positive punishment? Give an example of each.

A: Positive punishment adds an unpleasant stimulus to decrease behaviour (e.g., a speeding ticket). Negative punishment removes a pleasant stimulus to decrease behaviour (e.g., taking away a teenager's car keys for breaking curfew). Both reduce the target behaviour, but by different operations.


Connections to Other Topics

Learning connects directly to the memory chapter (Knowing Mind), since what is learned must be stored and retrieved. Classical conditioning reappears in discussions of phobias and anxiety disorders (fear conditioning) and in therapy (systematic desensitisation, aversion therapy). Operant conditioning principles underpin behavioural therapy and applied behaviour analysis. Observational learning connects to developmental psychology (children learning gender roles, aggression) and to social psychology (conformity, modelling prosocial behaviour).


Related Terms / Search Tags: three types of learning, classical conditioning Pavlov, operant conditioning Skinner, observational learning Bandura Bobo doll, positive reinforcement negative reinforcement, positive punishment negative punishment, reinforcement vs punishment, schedules of reinforcement fixed ratio variable ratio fixed interval variable interval, continuous vs partial reinforcement, extinction resistance, shaping successive approximations, token economy, behaviour modification, applied behaviour analysis, PSYCH 1100 final exam, Ohio State psychology