What a test asks
Standardised behaviour assessments give shelters repeatable data. Understanding what that data actually represents is the harder task.

What the protocol is doing03.1
A behaviour test is a sequence of controlled prompts delivered in a set order and scored against a rubric. The evaluator approaches the kennel, makes eye contact, presents a hand, introduces a food bowl, uses a doll or a gloved hand to interrupt feeding, tries a leash. Each step is recorded. The animal's responses — retreat, stillness, sniff, growl, snap — become a score, and the score becomes a decision path.
The logic is sound. Standardised protocols exist precisely because unaided human impressions are inconsistent: two experienced staff members watching the same dog will reach different conclusions depending on their history, their mood and which side of the kennel they stood on. A rubric removes some of that noise. The same prompt, delivered the same way, producing a scored response: that is a real improvement on pure gut feeling, and most shelters working at scale rely on some version of it.

What the rubric cannot do is control for the animal's condition on the day the test runs.
What the animal is experiencing03.2
An animal arriving at a shelter is frightened, under-slept, probably dehydrated, and navigating a building full of unfamiliar smells and sounds that do not resolve into anything it can predict. By the time a formal assessment is scheduled — often within the first few days — those baseline stressors have not cleared. They have sometimes intensified. A dog that has never growled over food in a home kitchen may guard a bowl fiercely in a kennel, because in a kennel, unpredictable things happen near food. A dog that froze on the test day might be sociable and relaxed a week later when it has learned that the morning routine is reliable.
The test captures a response. It does not capture a trait. Those are different things, and conflating them is where assessments cause harm.
The conditions inside a kennel block compound this problem. Sustained noise — concrete and mesh amplify it considerably — produces measurable changes in cortisol and in observable behaviour. An animal assessed in week one of shelter residence is being assessed at or near its worst: least acclimated, most aroused, often most reactive. The score sheet will not note this. The score sheet will record what happened when the evaluator raised the fake hand.
What the number means, and what it does not03.3
None of this means behaviour testing should be abandoned. A food-guarding response severe enough to break skin during assessment is real information, whatever its cause. A dog that shuts down entirely across multiple interactions over several days is telling you something. The issue is not the tool; it is the interpretation that follows.
A single-session assessment answers a narrow question: how did this animal respond to these specific stimuli, in this specific space, on this specific day? It does not answer how the animal will behave in a home, with a routine, with a person it has learned to trust — and the gap between those two questions is significant. What the test cannot tell you is, in practice, a long list.
The evaluator approaches the kennel, makes eye contact, presents a hand, introduces a food bowl, uses a doll or a gloved hand to interrupt feeding, tries a leash.
Good assessment practice treats a formal test as one data point among several: kennel observations logged over days, notes from the intake interview if the animal was surrendered rather than found, and — where possible — information from a foster placement. A foster carer seeing the animal in a domestic setting for two weeks will generate more usable behavioural information than any single protocol run in a kennel corridor. The test is a beginning, not a verdict.
The word "standardised" is worth sitting with. It means the prompts are consistent. It does not mean the animal's starting point is consistent, or that the kennel environment is neurologically neutral, or that the evaluator's read of "tense but non-reactive" matches the rubric's definition. Standardisation disciplines the input. It cannot discipline everything the animal carries into the room.

A test asks whether this animal, today, in these conditions, responded to this prompt in a way that looks like this category. That is a useful question. It is also a smaller question than it appears to be, and the score sheet will not remind you of the difference.