Pages

Showing posts with label behavior modification. Show all posts
Showing posts with label behavior modification. Show all posts

Sunday, July 10, 2016

The Nitty Gritty of Clean Training, Part 2: Reinforcement Schedules

As someone who uses reward based training, I hear this one all the time: "but if I train with treats, my dog will only ever listen when I have treats!"
That could be true- if you never change your reinforcement schedule and never fade out food treats. If you work with a trainer who understands reinforcement schedules and how to use them to fade out food treats and fine tune behavior, this is never a problem.

I will fight my natural tendency to give way more information than is necessary in this post, but this is one of those topics that gets a bit tricky so I use extra words to explain and re-explain myself. I apologize in advance for all the repetition.

A reinforcement schedule is a rule or pre-set program that determines how and when a response will be rewarded. Different stages of learning use different reinforcement schedules- when learning a new behavior we reward differently than when strengthening or proofing a known behavior. We also use different reinforcement schedules for different reasons in training, depending on the behavior we are encouraging or discouraging.

Look at you, learning about training! 
Easy, painless, and you are still awake. 
Let's dive in to the good stuff! 

When I being training a new behavior or cue with a dog, we will begin with a Continuous Reinforcement Schedule, which abbreviated as CRF. In this type of reinforcement, the dog gets a reward every time they offer the desired behavior. We use this when teaching new behaviors because we want the dog to learn that the new behavior is a really great thing to do- it always gets attention and a treat! If your dog gets a treat, pat on the head, and an enthusiastic "good boy!" every time he sits; he's going to try sitting more often. This comes in handy when we build behaviors on top of each other, because they always have a strong base behavior to fall back on if there is a regression in training. Regression can happen because training suddenly stops for a period of time or because of a change in environment/stimuli. When using a CRF, it is important to only use it until the dog understands the behavior, then switch to a less predictable schedule (this is also when we begin to give different rewards based on the quality of response, but that's covered below in Differential Reinforcement Schedules).

Partial Reinforcement Schedules (PRF) reward the desired response only after certain responses, either after a set ratio (number of responses) or interval (period of time). We can use these schedules to fine tune behavior once the basics are understood.
Within this schedule, there are five different types of reinforcement:
  1. Fixed Ratio: The dog gets a reward after a predetermined number of responses. For example, you can train your dog to "count" using this method by rewarding after say, three barks and labeling it "count to three". This could be done with any number, of course! 
  2. Variable Ratio: The dog gets rewarded after a different number of responses, but the average of them getting the reward is determined by you. If you want an average of three responses, you would reward for: 1, 4, 2, 3, 2. The average of these responses is 3. This is what I use to start fading out treats in training the 'heel'. At first, the dog is rewarded every step for staying in the 'heel' position. As the get better with staying in position in anticipation of treats, the treats are given after one step, two steps, four steps, three steps, two steps. They are getting rewarded on average every three steps, but it's not always three exactly and they are getting fewer treats than in initial training. Over time, we simply make the average a bigger number. 
  3. Random Ratio: This is the other way to build strong behaviors. In random ratio, the dog gets a reward sometimes, but not other times. It should be as random as possible. Truly random rewarding is hard for us people to wrap our heads around; we try to make patterns so it makes sense in our minds. Dogs are great at figuring out patterns, so they soon learn if we are actually making a pattern and predict it. This can be used in training the 'heel' just like the above example, but we would want to keep it random, instead of aiming for an average number of steps. 
  4. Fixed Interval: The dog gets a reward only when the behavior is offered when a set period of time has elapsed since the previous response. This is something that we don't really use much in training because it actually isn't terribly useful in most training situations. The idea is that a dog offers a behavior, like 'sit' and gets a reward. There would be a predetermined interval, lets say 4 seconds, that the dog needs to wait until it can offer the 'sit' again and get a reward. If they sit at 1, 2, or 3 seconds, they get no reward. Any 'sit' after 4 seconds gets a reward. Over time, responses on the part of the dog go up because they know they have to offer the behavior to get a reward. It's a fun thing to do, but really has little real use in day to day training. The problem is that a dog can get distracted and forget to respond with the correct behavior in that interval, so we can't effectively train anything that's well remembered.   
  5. Variable Interval: Just like variable ratio above, this is a reward for different responses, averaging a number you have picked. The difference is that this is rewarding for a period of time instead of a number of responses. The dog would get a reward for the correct behavior after a period of time has elapsed, but that interval of time will vary within an average. Like fixed interval this can result in a steady string of responses, but since the response is dependent on the animal offering it, can be tricky to use in training. 
When using a Differential Reinforcement Schedule, rewards are given after certain types of responses are offered or after certain rates of response are offered. Basically, this means that the dog gets a reward based on the quality of their response or the frequency of offering the correct response. This is what we use to fine tune behaviors, to build complex behaviors, or work with especially nervous, anxious, or reactive dogs.

 1. Response Type schedules are simply the quality of the response- a 'down' with the belly all the way on the ground is preferred over a 'down' with the belly tucked up and not touching the ground.

Within this, there are three types of schedules which we use to get the desired behavior and remove unwanted behaviors:
      a. Differential Reinforcement of Incompatible Behaviors (DRI): A dog jumps to greet people will be rewarded for any behavior that they can't do while jumping. Sitting, laying down, or simply standing would all be considered incompatible behaviors. These incompatible behaviors become more rewarding than the problem behavior (jumping).
      b. Differential Reinforcement of Other Behaviors (DRO): A dog who barks at passers-by on walks can be rewarded for doing anything that is not barking. These other behaviors become more rewarding than barking, so the barking starts to diminish. 
     c. Differential Reinforcement of Excellent Behaviors (DRE): A dog who perfectly heels on command when asked the first time, then sits in position when the handler comes to a stop would get a reward because that is an ideal response. We tend to reward these great responses a bit longer because they are the ultimate goal and we want them to become the normal. By rewarding these great behaviors, all others extinguish themselves. 

2. Response Rate Schedules are ones that require a dog to respond at a certain rate for that reward. The reward is based on the offering of the correct behavior within the correct time period. Much like fixed interval and variable interval training, these aren't of as much use in dog training, but here you go anyway. 
  • Differential Reinforcement of High Rates (DRH): A dog is only rewarded for offering the 'look' behavior if it occurs within 7 seconds of the previous response. If the dog looks at 1, 2, 3, 4, 5, 6, or 7 seconds, they get a reward. If it is 7 seconds or more they get no reward. This is used to build a steady stream of responses.
  • Differential Reinforcement of Low Rates (DRL): A dog is only rewarded for offering the 'look' behavior after a specified period of time has elapsed, lets say 7 seconds. Any look after 7 seconds gets a reward, anything before 7 seconds does not. 
A Duration Reinforcement Schedule requires the dog to respond throughout a set period of time; these periods of time can be fixed, variable, or random. The classic example for this is the 'stay' cue. A dog is asked to hold the stay position for a period of time. Initially in training, we work with a short period of time and build it up gradually to longer duration and out of sight stays.

  • Fixed Duration: The dog has to stay for 1 minute to get a reward. If they get up before that minute is up, there is no reward. 
  • Variable Duration: The dog has to stay for an average of 1 minute to get the reward. This is the best way to lengthen the duration of a stay because you are on average staying within the time period you know the dog will tolerate, but can gradually increase the duration by increasing the average.
  • Random Duration: The dog is asked to stay for random periods of time, rewarded only if they do so. This is a great way to lengthen duration also, because the dog can't predict how long you will be gone. If we simply leave for longer each time, the dog predicts that the time period will be longer, since they are good at putting together patterns. 

Still awake? Good job, you're almost done!

So what does all that mean? It means that you can fine tune and train different behaviors using different reinforcement schedules. Within this, you can even give different types of rewards based on responses (more on that another day).
There are three lessons I want you to come away with from this:
1. there is strong relationship between continuous reinforcement and degradation of behavior even before the food is faded.  If a behavior is always followed by a treat, over time the dog has no motivation to offer the behavior quickly or perfectly. If a behavior is always followed by a treat and the treats suddenly stop, the behavior stops too because the behavior is no longer paying off as it had been! Dogs who are on a continuous reinforcement schedule too long end up with sloppy or slow behaviors and behave like spoiled children, demanding things they want.
2.  Random and variable reinforcement always result in the strongest behaviors, with much lower incidents of the behavior extinguishing as rewards fade. If a behavior is always rewarded initially and then randomly or variably rewarded, there is still always the possibility of a reward, so the behavior continues with the same strength. This is how a slot machine works. The machines pay out on a variable or random schedule, though it is very difficult to predict exactly when it will. The longer you keep putting coins in, the more convinced you become that it will pay off next time.
3. It is very difficult for us humans to be truly random, which is why we tend to use variable rates of reinforcement in training. That way, your human need for some order is met and your dog is still not getting rewarded every single time, so we still get strong behaviors. 

The real point in telling you all of this, aside from giving you great reading material for your next bout of insomnia or a new drinking game (count how many times the word reinforcement is in here) is to demonstrate that the person who trains you and your dog should know a LOT about learning and training. It's not just a matter of tossing a collar on a dog and grabbing some treats- my 3 year old son can do that. It's not a matter of putting a pinch, choke, or prong collar on your dog and yanking him around to demonstrate "who is boss". Training and subsequent learning should be intentional, systematic, soundly based in science and well executed. There should be some room for flexibility with each individual dog/human pair and nobody should be pushed to the point of breaking or shutting down in training. Once you reach that point, nothing good is being taught.


Resources:
 Excel-Erated Learning; Explaining in Plain English How Dogs Learn and How Best to Teach Them by Pam Reid, pgs. 48-59

http://www.lifecircles-inc.com/Learningtheories/behaviorism/Skinner.html

http://www.educateautism.com/applied-behaviour-analysis/schedules-of-reinforcement.html

Tuesday, June 7, 2016

Conditioning: The Nitty Gritty of Clean Training, Part 1

Did you know that proper conditioning is important for your dog?

In training, we use both Operant and Classical Conditioning and I am using this post to tell you all about Operant Conditioning and hopefully not bore you too much. I'll go into Classical Conditioning in the next couple weeks, but for simplicity sake we will say for now that it's the part of training is learning by association.

Operant conditioning involves using reinforcers and punishers to get the desired behavior or stop an unwanted behavior.

Reinforcers-generally speaking, this is something the dog likes. It is important to keep in mind that reinforcers are not universal and therefore depend on the individual dog. Most dogs like food, so using treats in training will work for most dogs. Some dogs prefer a tennis ball, squeaky toy, belly rub, or playtime with another dog or their handler.
Reinforcers are used to encourage the repetition of a behavior. 
For example, a dog is asked to sit. If they sit, they get a tasty treat or a squeaky toy to play with for a minute.
Punishers- generally speaking, this is something that the dog doesn't like. Just like reinforcers, these also vary by individual dog. Some dogs don't like a stern verbal correction, some don't like being ignored or denied the opportunity to play. Most dogs don't like physical corrections because they are uncomfortable (or painful).
Punishers are used to decrease the repetition of a behavior. 
For example, a puppy starts biting their owner's hand during play. The owner can say "no" and walk away as a punishment. The puppy is losing the opportunity to play and has been given a verbal correction. 
I want to highlight again that there are many different types of punishers and many different types of punishers. In my experience, a lot of folks out there assume that a reinforcer is always food and a punisher is always pain. Since reinforcers vary by dog, how on earth could this be true?
I'll tell you a secret- it's not. I use both punishers and reinforcers in training: I don't limit myself to only rewarding with food and I steer clear of using physical corrections as punishers (we will get into why a little later). 

The next part of Operant Conditioning involves the application of these reinforcers and punishers and here is where it gets a little tricky. I have included a couple of great visual aids that I had nothing to do with creating so I'll credit them to where would up when I did a google search of the quadrants of Operant Conditioning. 
I'll start with the pretty pictures that I didn't put together:

This is from Fed Up Fred: 




This one is from a dog training forum, originally from a ClickerExpo: 



To use these reinforcers and punishers, we can give or take them away from our dog. Giving or adding a reinforcer/punisher is considered positive (+). Again, positive isn't necessarily a good thing, it simply means something is being added to the scenario as a result of the dog's behavior. It is being added to either encourage or discourage the behavior. 
Taking something away from the dog is negative (-). Negative isn't exclusively a bad thing, it just means we are taking away something from the situation or from the dog. It is being taken away to either encourage or discourage the dog's behavior. 
So, now we have reinforcers (R) and punishers (P); and positive (+), or negative (-) applications. 
Take a look at those charts again, or just look at the one you like best. 

R+ is Positive Reinforcement= something the dog likes is given to the dog to increase the behavior that immediately preceded it. The dog who gets a treat for sitting is getting positive reinforcement. 

R- is Negative Reinforcement= something the dog does not like is removed in an effort to increase the behavior that immediately preceded it. Pressure from a choke or prong collar is released once a dog stops pulling on leash. 

P+ is Positive Punishment= something the dog does not like is given to decrease the behavior that immediately preceded it. A "collar pop" is given as a response to a dog lunging on leash. 

P- is Negative Punishment= something the dog likes is removed in an effort to decrease the behavior that immediately preceded it. A dog jumps to greet me as I reach for a treat- I immediately put the treat away and turn to ignore the dog- he has (momentarily) lost the opportunity for treats and attention, which he likes. 

Now, remember earlier when I mentioned that people generalize reinforcement as treats and punishers as pain and how they are wrong in painting it all in black and white? Well, you can read that fourth paragraph again but I really did say it. I do use mostly R+ training, though I will use P- and R-. 
Here are a few examples that I have used just this week:
-A dog who runs to me quickly and immediately when I call him will get a treat and lots of praise and attention as a reward (to increase that behavior). This is R+
-A dog who jumps to greet or play will experience me walking away, putting my treats away or will get a 5 minute time out if he can't be redirected from the jumping. Since he is jumping for attention, I remove the thing he wants to decrease the jumping! This is P-
-A dog who is fearful of men in hats will get more distance from that scary guy in the hat 20 feet away if he can look at me or sit when asked. Something he doesn't like goes away when he offers the behavior I want, which is paying attention to and trusting me. This is R-

P+ is the one I really do avoid using, but that does not mean I don't understand how it works. I know that it is meant to stop behaviors quickly since the dog will be trying to avoid something they do not like. I know it can work or seem to work on plenty of dogs, my concern is more with the fallout from using such methods. Some dogs will react quite adversely to corrections like this and become aggressive or reactive in defense. This is because we can't actually say to a dog "ok, you will be getting a shock or "collar pop" now, because you ran after that kid on a bike".  For all we know, the dog just wants to run and play with the kid on a bike, but he may have different ideas and need to attack those scary tires. All our dog knows now is that whenever a bike goes by, something not so fun happens. This is where aggression and reactivity can increase because of P+. The other thing that I have seen happen is a dog actually shutting down and becoming fearful of bikes, children or their handler. If you want to really ruin your day, read about Learned Helplessness experiments that were done on dogs in the 1960's. You may be outraged about the fact that it happened 50 years ago, but what I agonize over is the fact there are trainers out there using very similar methods today in an effort to extinguish behaviors and train basic obedience skills. 
In addition, what usually happens isn't that a behavior is stopped- it's just suppressed. It may stay suppressed forever, or the dog can become like a ticking time bomb and one day they can't take it anymore. That's pretty extreme but I have seen it happen. I have seen playful, carefree puppies change to reactive, shut-down pups when a shock, prong, or choke collar is used. By studying canine body language, you can see for yourself that a dog who is being walked "under total control" is actually fearful and unsure. I don't know about you, but I'd rather my dog have a good time and be relaxed. So, yes P+ may work on dogs, but why take a chance on traumatizing your dog, ruining your bond and changing their personality? 

For more on why I don't like to use P+, check out my post on how a "calm, submissive" dog is an oxymoron. 

Next time, I will delve into using conditioning over time to ensure that your dog will respond in a variety of situations. 

Resources:

Excel-Erated Learning: Explaining in Plain English How Dogs Learn and How Best to Teach Them, By Pamela Reid


Don't Shoot the Dog: The New Art of Teaching and Training, By Karen Pryor 

http://www.simplypsychology.org/operant-conditioning.html

http://www.britannica.com/topic/learned-helplessness