Variable Reinforcement is a LIE
Ivan BalabanovIvan Balabanov · 영상 · 21분
가변 강화는 거짓말입니다
Variable Reinforcement is a LIE
0:00
자, 만약 당신이 간헐적 강화 계획에 대해 들어온 모든 내용이 절반의 진실에 불과하다면 어떨까요? 오늘은 제가 수강생들과 다루는 주제들을 맛보기로 보여드리고, 여러분이 생각해 볼 만한 거리를 드리고 싶습니다. 강화 계획은 반려견 훈련과 심리학 분야에서 매우 흔하게 언급되는 개념입니다. 분명히 훈련사들이 '간헐적 강화 계획으로 바꾸세요, 무작위 보상이 행동을 더 강화해주니까요'라고 말하는 것을 들어보셨을 겁니다. 그러고 나면 그들은 슬롯머신 예시를 들겠죠. 사실 여러분 중 많은 분도 그렇게 하고 계실 겁니다. 물론 그 말에도 일리는 있습니다. 하지만 거기서 멈춘다면, 아마 더 큰 그림을 놓치고 계신 걸지도 모릅니다. 제가 훈련에서 강화 계획을 사용하는 방식은 슬롯머신 논리를 살짝 넘어섭니다. 그럼 간단히 설명해 보겠습니다. 하지만 먼저 기본부터 시작하죠. 행동 심리학에는 여러 가지 강화 계획이 존재합니다. 연속 강화가 있는데, 기본적으로 모든 올바른 반응마다 보상을 주는 방식입니다. 항상 듣는 조언이자 사실이기도 한데, 새로운 행동을 가르칠 때 가장 좋은 방법입니다. 고정 비율 계획도 있습니다. 이는 정해진 횟수의 반응 후에 보상을 제공하는 방식입니다. 예를 들어, IGP 훈련의 '짖고 버티기(bark and hold)'를 봅시다. 훈련사가 세 번 짖을 때마다 보상을 준다면, 반려견은 계속 짖기보다는 세 번 짖고 보상을 기다릴 가능성이 큽니다. 그게 바로 고정 비율 계획의 예시입니다. 그리고 변동 비율 계획이 있는데, 무작위 강화 계획이라고도 불립니다. 이는 예상치 못한 횟수의 반응 후에 보상을 주는 방식입니다. 여기서 라스베이거스 비유가 등장하는 것이죠. 때로는 두 번 후에, 때로는 다섯 번 후에 보상을 주는 식입니다. 또 다른 하나는 고정 간격 계획입니다. 일정한 시간이 지난 후 첫 번째 반응에 대해 보상합니다. 예를 들어, 개가 30초가 지난 후에 앉았을 때만 보상을 받는 식이죠. 변동 간격 계획입니다. 변동 간격 계획이요. 예측할 수 없는 시간 간격이 지난 후의 첫 번째 반응에 대해 보상합니다. 어떨 때는 10초 후가 될 수도 있고, 또 어떨 때는 1분 후가 될 수도 있죠. 이게 강화 계획의 거의 전부라고 보시면 됩니다. 사실 반려견 훈련 현장에서는 우리가 이 대부분을 의도적으로 사용하지는 않습니다. 훈련사들은 보통 기술을 가르칠 때는 연속 강화 방식을 고수하고, 행동을 유지하려 할 때는 변동 강화 방식을 사용하죠. 바로 여기서 슬롯머신 비유가 항상 등장하죠, 그렇죠?
All right, what if everything you've been told about variable reinforcement schedule is only half the story? Today I want to give you a taste of the kind of topics I go over with my students and hopefully give you something to think about. Schedule of reinforcement is something that gets thrown around a lot in dog training and in psychology. I'm sure you've heard trainers say switch to variable reinforcement schedule because random rewards make the behavior stick. And then they're going to give you the slot machine example. In fact, probably many of you are doing the same. And sure, there is truth to that. But if you stop there, you are probably missing the bigger picture. The way I use reinforcement schedules in training goes a little bit beyond slot machine logic. So let me break this down quickly. But first let's start with the basics. In behavior of psychology, there are several reinforcement schedules. We have continuous reinforcement, which basically every single correct response is rewarded. And the advice that's always given, and it's true, is that it's the best option for teaching new behaviors. We also have fixed ratio. That's where the reward is offered after set number of responses. For example, like in, let's say, IGP, the bark and hold. If the trainer rewards every third bark, then the dog most likely is going to bark three times and wait for the reward instead of continuous barking. That's your fixed ratio example. And we have variable ratio, also known as the random schedule of reinforcement. This is where reward, we reward after unexpected number of responses. That's where the Las Vegas analogies come. Sometimes we pay after two, sometimes after five responses, and so on. Another one, fixed interval. We reward for the first response after a set period of time. For example, the dog only gets reinforced if it sits after 30 seconds have passed. Variable interval. Variable interval. We reward for the first response after unpredictable time intervals. So sometimes it can be after 10 seconds, sometimes after one minute. That's pretty much the full manual of reinforcement schedules. Now the reality in dog training, we don't use most of this deliberately. Trainers typically stick with the continuous reinforcement when teaching and variable reinforcement when trying to maintain the behavior. That's where the slot machine analogy always comes in, right?
3:59
하지만 주의해야 할 다른 점도 있습니다. 가끔 훈련사들은 자신도 모르게 고정 패턴이나 간격 패턴에 빠지곤 합니다. 예를 들어, 항상 힐링(나란히 걷기) 세 걸음 후에 보상하거나, 1분간의 엎드려 기다리기 훈련이 끝날 때마다 보상하는 식이죠. 훈련사가 의도하지 않았더라도 개는 그런 패턴을 알아차립니다. 그러니까 이론적으로는 이런 모든 강화 계획이 존재하지만, 실제 반려견 훈련에서는 훈련사들이 연속 강화와 변동 강화 계획에 크게 의존합니다. 그리고 사실 대부분의 논의가 바로 여기서 막히게 되죠. 많은 훈련사가 처음에는 연속 강화로 시작했다가 개가 행동을 익히고 나면 변동 강화로 전환하는 이유가 바로 이것입니다. 물론 여기에는 과학적 근거가 있습니다. 스키너도 있고, 소거 저항 연구 같은 것들이 전부 다 관련되어 있죠. 맞아요, 변동 강화 계획으로 훈련된 행동에 대해 보상을 완전히 중단하면, 연속 강화로만 훈련된 경우보다 보통 더 오래 지속됩니다. 그러니 이론적으로는 완전히 말이 되죠. 하지만 그 슬롯머신 예시를 더 자세히 들여다봅시다. 여기 우리가 간과하거나 미처 생각하지 못하는 부분이 있습니다. 대부분의 사람들은 슬롯머신에 중독되지 않습니다. 수백만 명이 라스베이거스를 방문해 재미로 조금씩 즐기다가 자리를 뜹니다. 제가 바로 그 대표적인 예입니다. 오직 특정 유형의 사람만이 도박 중독의 굴레에 빠지게 됩니다. 왜 그럴까요? 가변 강화 계획이 마법처럼 모두를 중독시키는 것은 아니기 때문입니다. 중독은 그 계획이 개인과 상호작용할 때 발생합니다. 즉 그 사람의 성격, 뇌 화학 작용, 취약성 등이 복합적으로 작용하는 것이죠. 개들도 이와 다르지 않습니다. 모든 개가 예측 불가능성에 의해 동기 부여되는 것은 아닙니다. 많은 경우, 가변 강화 계획이 반드시 집착을 만들어내는 것은 아닙니다. 그저 하나의 패턴일 뿐입니다. 따라서 슬롯머신 비유는 다소 지나친 단순화입니다. 지속성은 단순히 무작위성만의 문제가 아닙니다. 정말로 중요한 것은 행동 자체가 즐겁고 의미 있으며 습관이 될 때입니다. 그것이 바로 행동을 탄탄하게 만드는 핵심입니다. 하지만 여기서 진짜 문제는 따로 있습니다. 개들은 스키너 상자 속의 비둘기가 아니며, 슬롯머신 레버를 당기는 도박꾼들도 결코 아닙니다. 그래서 제게 그런 식의 논리는 실험실 수준에서 멈추는 것입니다. 훈련에서 행동이 유지되는 것은 단순히 무작위성 때문만이 아닙니다. 그 행동이 습관이 되었기 때문에 유지되는 것이며, 개 자신이 그 행동을 진심으로 즐기기 때문에 유지되는 것입니다. 일단 행동이 진정한 습관이 되면, 소거는 더 이상 위협이 되지 않습니다. 그리고 행동이 자기 강화적일 때, 개가 그 작업 자체에서 기쁨을 느낄 때, 어떤 슬롯머신도 그것을 이길 수 없습니다. 제 훈련 방식에서 두 가지 짧은 예시를 들어보겠습니다. 힐링(따라 걷기)을 예로 들어보죠. 힐링을 가르치기 시작할 때, 물론 저는 연속 강화 계획을 사용합니다. 따라서 올바른 동작마다 보상을 줍니다.
But there is also something else to watch for. Sometimes trainers fall into fixed or interval patterns without ever realizing it. Maybe they always reward after exactly three steps of healing, or always at the end of one minute down stay. The dog notices those patterns, even if the trainer doesn't intend them. So, yes, technically all the schedule exists, but in dog training, in practice, dog trainers lean heavily on continuous and variable reinforcement schedule. And that's where most of the conversation really gets stuck. This is why many trainers start on continuous reinforcement and then switch to variable once the dog knows it. Sure, there is science behind this. You have Skinner, you have resistance to extinction research, and the whole thing. And yes, if you stop rewarding completely a behavior that is trained on variable reinforcement schedule, it will usually last longer than one trained only on continuous. So, in theory, it makes total sense. But let's look closer at that slot machine example. Here is something we kind of forget or don't think of. Most people are not hooked on slot machines. Millions go to Vegas, play a little for fun, and walk away. I'm a prime example of this. Only a certain type of person gets trapped in the cycle of gambling addiction. Why? Because variable schedules don't magically hook everyone. Addiction happens when the schedule interacts with the individual, their personality, the brain chemistry, susceptibility, and so on. And dogs are not different. Not every dog is motivated by unpredictability. For many, variable schedule of reinforcement doesn't necessarily create obsession. It's just a pattern. So, the slot machine analogy is kind of oversimplification. Persistent isn't just about randomness. What really matters is when the behavior itself becomes enjoyable, meaningful, and habitual. That's what makes it bulletproof. But here is the real problem. Dogs aren't pigeons in Skinner boxes, and they're definitely not gamblers pulling levers at the slot machine. So, that kind of logic stops in the laboratory to me. In training, behaviors don't just survive because of randomness. They survive because they become habits, and because the dog actually loves doing them. Once a behavior is a true habit, extinction isn't really a threat. And when a behavior is self-reinforcing, when the dog takes joy in the work itself, no slot machine can compete with that. I'll give you two quick examples from my own training. Let's take healing. When I start teaching healing, of course, I use continuous schedule of reinforcement, so every correct step gets paid.
7:58
그 명확성이 필수적이죠, 그렇죠? 하지만 시간이 지나면 무언가 바뀝니다. 제 개들에게 힐링(healing)은 단순한 대가 이상의 의미가 됩니다. 그것은 습관이자 즐거움이 됩니다. 개들은 도전을 갈망하고, 정밀함과 저와의 상호작용을 원하게 되죠. 그 시점이 되면 간헐적 강화 계획이 힐링을 유지하는 핵심 요소가 아닙니다. 실제로는 개가 느끼는 즐거움 자체가 핵심이죠. 이제 센드 어웨이(send away)를 예로 들어보죠. 더 먼 거리, 더 빠른 속도, 더 날카로운 정밀함을 원한다면, 그때도 무작위 보상 계획인 간헐적 강화로는 목표를 달성할 수 없습니다. 이럴 때 연속 강화 계획이 다시 빛을 발합니다. 모든 개선과 더 날카로운 노력 하나하나에 보상하죠. 그것이 바로 행동의 기준을 높이는 방법입니다. 이렇게 생각하시면 됩니다. 간헐적 강화 계획은 행동을 현재 수준에서 고착화하는 경향이 있습니다. 연속 강화 계획은 행동을 더 정교하게 다듬게 해주죠. 자, 이제 일반적인 강화 계획에 대한 가르침과 완전히 반대되는 이야기를 하겠습니다. 대부분의 훈련사들과 교재들은, 이렇게 말할 것입니다. 행동을 가르치는 습득 단계에서는 연속 강화 계획을 사용하고, 그다음 빠르게 간헐적 강화 계획으로 넘어가서 행동을 더욱 지속 가능하게 만들고 소거에 저항하도록 만들라고 말이죠. 유지 단계에 대해서도 물론, 그런 이유로 간헐적 강화 계획을 계속 유지하라고 조언할 것입니다. 그것이 통념이죠. 하지만 저는 더 효과적인 방법을 찾았습니다. 저는 종종 연속 강화 계획을 유지 단계에서도 계속 사용합니다. 그리고 저는 그것을 이른바 높은 수준의 유지가 필요한 행동들에 적용합니다. 무슨 뜻일까요? 앉아를 예로 들어보죠. 대부분의 개들에게, 앉기는 자연스럽고 쉬운 반응이며, 배우기 쉬운 동작입니다. 일단 훈련이 되면, 매번 적절한 수준의 앉기를 보여줄 겁니다. 하지만 대회에서는, 적절한 수준의 앉기만으로는 충분하지 않습니다. 우리는 더 날카롭고, 더 빠르며, 더 정교한 것, 개의 자연스러운 행동보다 더 나은 무언가를 원하며, 저는 그것을 이렇게 해냅니다. 그 높은 수준을 끌어내기 위해, 저는 계속해서 보상을 줍니다. 연속 강화 계획을 통해서요. 왜냐고요? 제가 개의 기본 행동보다 더 많은 것을 요구할 때는, 그 날카로움을 유지하기 위해 끊임없는 강화가 필요하기 때문입니다. 물론, 거기에 따르는 주의 사항이 있습니다. 약간 더 높은 수준까지만 요구할 수 있다는 점이죠. 만약 너무 무리하게 밀어붙여서, 계속해서 더 많은 것을 요구한다면, 결국 한계점에 다다르게 될 것입니다. 그러면 행동이 개선되는 대신, 무너지기 시작할 겁니다. 왜냐하면 개는 말 그대로, 자신의 한계치보다 더 잘 수행할 수는 없기 때문입니다. 여기 또 다른 요점이 있습니다. 훈련사들은 이렇게 말하곤 하죠, 더 나은 시도들에 대해서만 간헐적 강화 계획으로 보상하라고요. 그러면 기준을 높일 수 있게 될 것입니다. 하지만 실제로는 그렇지 않습니다.
That clarity is essential, right? But over time, something changes. For my dogs, healing becomes more than just a paycheck. It becomes a habit and a joy. They crave the challenge, the precision, the interaction with me. At that point, variable schedule of reinforcement isn't what holds healing together. It is actually the dog's own enjoyment. Now, let's take the send away. If I want more distance, more speed, sharper precision, then, again, random reinforcement schedule is not going to get me there. This is where continuous schedule of reinforcement shines again. I reward every improvement, every sharper effort. That's how I will raise the criteria for the behavior. So, you can think of this way. Variable schedule of reinforcement tends to freeze the behavior where it is. Continuous schedule lets me sharpen it. Now, here is something that goes completely against the way reinforcement schedules are usually taught. Most trainers, and most textbooks, if you want, will tell you, use continuous schedule of reinforcement during the acquisition state to teach the behavior, then quickly move to variable schedule of reinforcement to make the behavior more persistent and resistant to extinction. And for the maintenance stage, of course, they would advise to stay with variable schedule of reinforcement because of that. That's your conventional wisdom. But here is what I found works better. I often continue to use continuous schedule of reinforcement during the maintenance stage. And I do it for what I call high-maintenance behaviors. What do I mean? Let's take the sit as an example. For most dogs, sitting is a natural and easy response, easy thing to learn. Once trained, we'll get a decent sit every time. But in competition, a decent sit isn't enough. We want something sharper, something faster, more precise, something better than the dog's natural, this is how I do it. And so to get that extra level, I keep rewarding on continuous schedule of reinforcement. Why? Because when I am asking for more than the dog's default behavior, it needs constant reinforcement to maintain that sharpness. Now, of course, there is a warning that goes with that. Like you can ask only for a little more. If we push too far, if we demand more and more and more, eventually we're going to hit a breaking point. And instead of improving the behavior, it's going to start to crumble because the dog literally cannot perform better than its limit. And here is another point. Trainers will say, just reward the better attempts on a variable schedule and you will be able to raise the criteria. But it doesn't work that way.
11:58
만약 개가 이미 변동 강화 스케줄에 있다면, 보상을 주지 않는 것은 개에게 전혀 특별한 일이 아닙니다. 개는 실망해서 더 열심히 하거나 더 자주 시도하지 않을 것입니다. 그저 아마 다음번에는 평소처럼 보상을 받겠지라고 생각할 뿐입니다. 그래서 저는 유지 관리가 많이 필요한 행동에 대해서는 연속 강화 스케줄을 유지합니다. 이해가 되셨기를 바랍니다. 다시 말해, 무작위 스케줄은 행동을 유지하게 만든다고 할 수 있습니다. 하지만 행동을 더 좋게 만드는 것은 바로 연속 강화 스케줄입니다. 자, 이제 사람들이 잘 이야기하지 않는 부분에 대해 말씀드리겠습니다. 저에게는 이것이 정말 고급 훈련의 핵심입니다. 우리는 보통 소거 폭발을 부정적인 방식으로 듣습니다. 개가 문 앞에서 짖을 때, 우리가 보상을 중단하면, 한동안은 짖는 행동이 더 심해졌다가 점차 사라지게 됩니다. 그래서 훈련사는 이렇게 말할 것입니다. 네, 그게 전형적인 소거 폭발이니, 그냥 대비하세요. 하지만 소거 폭발은 단지 장애물만은 아닙니다. 소거 폭발은 행동을 개선하기 위한 가장 강력한 도구 중 하나가 될 수 있습니다. 그 방법은 다음과 같습니다. 제가 연속 강화 스케줄을 사용하고 있을 때, 제 개는 매번 보상을 기대합니다. 그래서 제가 갑자기 보상을 주지 않으면, 개는 어떻게 행동할까요? 항의. 이봐, 내 돈은 어디 있지? 그리고 그 항의 과정에서, 개는 행동을 과장합니다. 그 과장된 행동은 아주 귀중하죠. 그렇게 저는 힐(heel)을 날카롭게 다듬고, 웨이스트(waist)를 고정하고, 앉기 등을 완성합니다. 소거 폭발(extinction burst)은 개가 더 많은 행동을 하도록 유도하고, 저는 그때 보상을 줍니다. 그것이 제가 기준을 높이는 방법입니다. 자, 이미 변동 강화 계획을 따르고 있는 개에게 이 방법을 시도한다고 상상해 보세요. 효과가 없을 겁니다. 개는 항의하지 않거든요. 그냥 이렇게 생각하죠, 좋아, 다음번엔 보상을 받겠지. 폭발(birth)이 일어날 이유가 없는 겁니다. 추가적인 노력을 기울일 이유가 없죠. 그러니까 네, 변동 강화가 도움이 되고, 행동을 유지하게 해주기는 합니다. 하지만 저에게 행동을 더 나은 수준으로 만들 수 있는 레버리지를 주는 것은 연속 강화 계획입니다. 그것은 많은 훈련사가 완전히 간과하는 엄청난 차이입니다. 마무리하기 전에, 훈련 중에 나타나는 관련 내용을 잠시 언급하겠습니다. 가끔, 개들이 짖거나, 낑낑대거나, 심지어 핸들러를 살짝 깨물기도 합니다, 수행 중에 말이죠. 많은 이들은 이것을 리킹(leaking)이라고 부릅니다. 물론 어떤 이들은, 드라이브(drive)라고 부르기도 하죠. 또 다른 이들은 이것을 불복종과 혼동합니다. 하지만, 대부분의 경우, 이는 두 가지 요인으로 귀결됩니다. 명확성 부족 무엇을 기대하는지에 대한 또는 잘못된 다루기 강화 스케줄. 개는 알고 있다면 무언가를 해야 한다는 요구가 있다는 것을, 하지만 무엇을 해야 할지 완전히 이해하지 못한다면, 그 좌절감은, 당연하게도, 밖으로 표출됩니다. 만약 강화가 잘못 다뤄졌다면, 예를 들어, 지속적인 강화 스케줄에서 빠져나올 때 너무 빨리 진행했다면요. 그 항의는 나타납니다 소리를 내거나 무는 것으로, 방금 말했듯이요. 자, 여기가 핵심입니다.
If the dog is already on variable reinforcement schedule, withholding the reward is nothing unusual to them. The dog isn't going to be disappointed and try harder or often more. It just assumes maybe next time I will get paid as I always do. That's why I stay with a continuous schedule for reinforcement for high maintenance behaviors. Hope that makes sense. So, again, random schedules, we can say, keeps things alive. But it is the continuous schedule of reinforcement that makes behaviors better. Now, here is the part that not many talk about. And for me, it's really the essence of advanced training. We usually hear about extinction bursts in a negative way. A dog barks at the door, we stop reinforcing it, and for a while, the barking gets worse before it fades out. So, a trainer is going to say, yeah, that's a typical extinction burst, so just be ready for it. However, extinction bursts aren't just an obstacle. They can be one of the most powerful tools for improving behavior. Here is how. When I'm using a continuous reinforcement schedule, my dog expects a reward every single time. So, if I suddenly hold back, the dog is going to what? Protest. Hey, where is my money? And in that protest, the dog exaggerates the behavior. That exaggeration is gold. That's how I sharpen the heel, fasten the waist, sit, and so on. The extinction burst pushes the dog to offer more, and then I reward that. That's how I raise the criteria. Now, picture trying this on a dog already on variable schedule of reinforcement. It's not going to work. The dog doesn't protest. It just assumes, okay, maybe I get paid next time. There is no reason for birth. There is no reason to make an extra effort. So, yes, variable reinforcement helps, keeps the behavior alive. But it's the continuous reinforcement schedule that gives me the leverage to make behaviors better. And that's a huge difference most trainers completely miss. Before I wrap up, let me touch on something related that shows up in training. Sometimes, dogs will bark, whine, or even nip at the handler during performances. Many call this leaking. Some, of course, call it drive. Others confuse it with disobedience. But, most of the time, it comes down to two things. Lack of clarity about what's expected or mishandling reinforcement schedules. If the dog knows there is a demand to do something, but doesn't fully understand what, the frustration, of course, leaks out. If reinforcement has been mishandled, let's say, coming off continuous schedule of reinforcement too quickly. The protest shows up as vocalizing or biting, as I just said. So, here is the key.
15:59
그 항의는 항상 나쁜 것은 아닙니다. 소거 폭발처럼, 양날의 검과 같습니다. 당신이 그것을 인식한다면, 그것을 이용할 수 있습니다. 그 짖음, 살짝 무는 것, 작은 폭발과 같은 좌절감은 더 날카로운 노력으로, 더 빠른 속도나 더 많은 강도로 바꿀 수 있습니다, 당신이 무엇을 추구하든 말이죠. 하지만 안타깝게도, 모든 사람이 이렇게 보는 것은 아닙니다. 소셜 미디어 인플루언서 중 한 명이 떠오르는데 항상 추천하곤 하죠 개에게 도미넌트 칼라를 사용해 숨을 쉬지 못하게 하라고요 개가 소리를 낼 때마다 말이죠. 그리고 슬픈 점은 트레이너들이 그렇다는 것입니다 좋아요 버튼을 누르며 이것 참 훌륭하네라고 생각하고 계실 겁니다. 그렇지 않습니다. 그건 형편없는 조언이니까요. 왜냐하면 훈련 중에 짖거나 낑낑거리는 행동은 강아지의 목을 조른다고 교정되는 것이 아니기 때문입니다. 명확함과 숙련된 강화 핸들링으로 교정되는 것이죠. 강아지는 나쁜 행동을 하는 게 아니라, 당신에게 피드백을 주는 겁니다. 그 피드백을 제대로 읽어낸다면, 갈등 대신 훌륭한 결과로 바꿀 수 있습니다. 더 큰 그림을 보여드리겠습니다. 가변 강화 계획은, 다시 말하지만, 매우 유용하지만, 만능 해결책은 아닙니다. 연속 강화 계획은 기준을 높이고 정확도를 향상하는 데 결정적입니다. 소거로 인한 폭발은 두려워할 대상이 아닙니다. 그것은 행동을 더 명확하게 다듬기 위한 최고의 도구 중 하나입니다. 짖음, 낑낑거림, 입질, 등의 행동은 종종 강화 과정에서 발생하는 갈등으로 인한 항의이며, 이것은 충분히 해결될 수 있습니다. 생산적으로 유도된 만약 당신이 무엇을 하는지 안다면 말이죠. 라스베이거스 비유에 관해서는, 네, 앞서 말했듯이, 그건 이야기의 절반일 뿐입니다. 끈기는 단순히 무작위성에 관한 것이 아닙니다. 진정한 지속성은 상호작용 자체의 습관과 즐거움에서 비롯됩니다. 목표는 행동을 우연에 기대어 유지하는 것이 아닙니다. 그것은 행동을 매우 명확하고, 매우 강력하며, 매우 즐겁게 만들어 습관이 되게 하고, 자기 강화적이게 하여, 결과적으로 완벽하게 만드는 것입니다. 반려견을 훈련하고 있다면, 보상을 무작위로 주는 것에 집착하지 마세요. 스스로에게 물어보세요, 내가 연속 강화 계획을 진행하는 동안 행동을 아주 명확하게 만들었는가? 강화. 내가 강화를 단순히 유지하는 것이 아니라 품질을 높이는 데 사용하고 있는가? 우리 개가 실제로 훈련을 즐기고 있는가? 물론입니다. 물론이죠. 당연한 말입니다. 그럼요. 그렇고말고요. 그것이 즐거워하나요 상호작용을 충분히 그래서 습관이 되나요? 만약 당신이 그곳에 집중한다면, 소거는 당신이 생각하는 그런 문제가 아닙니다. 훈련사들이 그렇게 말하는 것만큼요. 그래서, 이것은 단지 많은 것들 중 하나일 뿐입니다 흥미로운 것들 중에서 제가 가르치는 제 학교에서 반려견 훈련사들을 위한, 훈련 과정이죠, 갈등 없이 하는 훈련 말입니다. 물론, 우리는 다룹니다 강화 계획들을, 하지만 우리는 교과서를 넘어섭니다 교과서를 넘어, 그리고 제가 부르는 것 이상의 뻔한 내용들 너머를요. 우리는 가르칩니다 어떻게 사용하는지 소거 폭발을 하나의 도구로, 어떻게 기준을 높이는지, 어떻게 해석하는지 발성 및 갈등 행동들을, 그리고 어떻게 만드는지 행동들을 스스로 강화되는 그런 행동들로 말이죠. 왜냐하면 진정한 훈련은 단지 슬롯 머신에 관한 것이 아닙니다. 그것은 명확함, 즐거움, 정확성,
That protest isn't always bad. Just like extinction bursts, it's a two-edged sword. If you recognize it, you can harness it. That bark, the nip, the little explosion of frustration can be channeled into sharper effort, more speed or more intensity, whatever you're after. But unfortunately, not everyone sees it this way. There is a social media influencer that comes to mind who always recommends stopping dog's air supply with a dominant collar every time it vocalizes. and the sad part is trainers are hitting the like button and thinking this is brilliant. It's not. It's a horrible advice because barking or whining in training isn't fixed by choking out the dog. It's fixed by clarity and skillful reinforcement handling. The dog isn't misbehaving, it's giving you feedback and if you read that feedback properly, you can turn it into brilliance instead of a conflict. Here is the bigger picture. Variable schedule, once again, very useful, but they are not the magic bullet. Continuous schedule of reinforcement is critical for raising criteria and improving precision. Extinction births aren't just something to fear. They are one of the best tools for sharpening behavior. The barking, whining, nipping, and so on are often protests born of reinforcement conflict which can be channeled productively if you know what you are doing. As far as the Las Vegas comparison, yeah, as I said, it's only half the story. Persistence isn't just about randomness. Real durability comes from habit and joy of the interaction itself. The goal isn't to keep behaviors alive by chance. It's to make them so clear, so strong, and so enjoyable that they become a habit, self-reinforcing, and therefore bulletproof. If you're training your dog, don't obsess over making rewards random. Ask yourself, have I made the behavior crystal clear during the continuous schedule of reinforcement? reinforcement. Am I using reinforcement to raise quality, not just keep things going? Does my dog actually enjoy the work? Of course. Does it enjoy the interaction enough that it becomes a habit? if you focus there, extinction isn't the problem trainers make it out to be. So, this is just one of the many interesting things that I teach at my school for dog trainers, training without conflict. Of course, we cover reinforcement schedules, but we do go beyond the textbook and beyond what I call obvious content. We teach how to use extinction bursts as a tool, how to raise criteria, how to interpret vocalization and conflict behaviors, and how to create actions that become self-reinforcing. Because real training isn't only about slot machines. It's about clarity, joy, precision,
19:59
교감에 관한 것입니다. 청취해주셔서 감사합니다. 어떻게 생각하시는지 알려주세요. 감사합니다.
connection. Thanks for listening. Let me know what you think. Thank you.