When Teddy Bears Do “Neurofeedback”

A satirical 2011 study used stuffed animals as “participants” to test whether NeurOptimal’s divergence index reflects real brain activity or just a clever algorithm.


If you’ve ever wondered whether a neurofeedback system could “improve” a teddy bear’s brain, a 2011 scoping study by Rubén Pérez‑Elvira and María Gracia Carpena‑Niño did exactly that… and not entirely as a joke. Working with the NeurOptimal system by Zengar, the authors strapped electrodes onto three stuffed animals and treated them as if they were human participants, complete with assigned ages spanning adolescence, young adulthood, and old age. Their aim was simple but sharp: if an automated neurofeedback system shows robust training effects in plush toys, how confidently can we interpret similar changes in actual patients?

This article comes from the archives, but its message feels surprisingly current, and maybe even more relevant now that “effortless brain optimization” tools are all over the marketplace. Let’s be clear from the outset: they might borrow the name, but these things are NOT neurofeedback.

Neurofeedback, in general terms, is a form of biofeedback where we measure brain activity—usually via EEG—and feed that information back to the person in real time so they can learn to regulate it. Biofeedback more broadly uses signals like heart rate, muscle tension, or skin conductance to support self‑regulation. At its best, neurofeedback is a sophisticated learning tool, helping brains adjust patterns linked to attention, mood, sleep, or arousal.

But what happens when the learning system is a sealed black box, and we aren’t entirely sure what the feedback metric represents, or even whether it reflects the brain at all? NeurOptimal, which is marketed as a form of neurofeedback, does not use qEEG‑guided assessment, does not let the clinician target specific EEG bands or regions, and does not provide transparent training parameters. Instead, it relies on a proprietary algorithm (a “divergence” index) that supposedly reflects how “upset” the brain’s electrical functioning is, with values trending toward zero as the system “optimizes” activity.

The authors of this study decided to stress‑test that assumption in the most charmingly absurd way possible: by treating stuffed animals as clients and seeing whether they “improve” with training.


Methods

The design of this study reads like a crossover between a methods section and a sketch from a neuroscience comedy special, but underneath the satire is a serious logic. Three stuffed animals were selected from a group of six, chosen for the similarity of their material composition. These plush toys were then assigned to three age categories, forming a 3×3 matrix of “participants”:

  • Group A: “teenagers” (ages 10, 14, 16)

  • Group B: “young adults” (ages 40, 40, 44)

  • Group C: “elderly” (ages 74, 80, 82)

Of course, the toys themselves did not age between sessions; instead, the researchers used the NeurOptimal user interface to assign different ages and IDs to the same physical stuffed animals, effectively creating nine “participants” from three pieces of inert matter.

Each “participant” received sessions with NeurOptimal according to the manufacturer’s recommendations: a pre‑training baseline, a period of neurofeedback training, and a post‑training baseline. Electrodes were attached using Ten20 conductive paste, and to ensure consistent placement across sessions, the authors left visible “teddy bear marks” of dried paste on the fur as a guide. This detail, described in the discussion, is both endearing and methodologically sound: the same non‑brain location on each toy received electrodes every time, reducing variability due to placement differences.

To further limit bias, each stuffed animal underwent three sessions in sequence—one for each age condition—without removing the electrodes. In practice, this meant that “Participant 1A” (teen), “Participant 1B” (young adult), and “Participant 1C” (elderly) were all the same toy, measured three times in a row while only the age data in the software changed. Raw data from the NeurOptimal system were exported as the proprietary “divergence index” for pre‑ and post‑training baselines. Higher divergence values were interpreted by the software as more dysregulated electrical activity; values closer to zero represented supposedly more “optimal” functioning.

The researchers entered these divergence scores into Excel for preliminary analysis and planned further work using SPSS. They computed mean pre‑ and post‑training values for each of the nine “participants,” as well as group averages for adolescents, young adults, and elderly. They then looked at correlations between age and divergence, and between the number of sessions and changes in divergence.

In other words, they treated this as a legitimate neurofeedback outcome study—except the “brains” were stuffed with foam, not neurons.


Results

Even though the subjects were inanimate, the data behaved as if real physiological change were occurring. Across all nine “participants,” the divergence index decreased from pre‑ to post‑training baselines. The bar chart on page 3 of the article (Progression by subjects) shows a consistent pattern: for every stuffed animal in every age condition, the pre‑training bars (blue) are higher than the post‑training bars (purple), suggesting that the system perceived each toy as becoming more “regulated” over the course of training.

When the authors averaged scores across groups, the same pattern persisted. Table 1 and the chart on page 4 illustrate that group means decreased substantially from pre to post for adolescents (Group A), young adults (Group B), and elderly (Group C). The total average divergence across all sessions dropped from approximately 217.6 pre‑training to 102.8 post‑training—almost a halving of “dysregulation” in the plush population.

The correlations add another layer of dark comedy. Age correlated positively with initial baseline divergence (ρ ≈ 0.61), meaning that “older” stuffed animals started with higher divergence scores, while younger ones looked more “electrically healthy.” When the authors aggregated data by group means, this relationship became extremely strong (ρ ≈ 0.99 at pre‑baseline), suggesting that the system’s divergence index tracked the user‑entered age field remarkably well—even though the underlying physical object never changed.

After training, age was only weakly correlated with post‑baseline divergence at the individual level (ρ ≈ 0.14), but still showed a strong positive relationship at the group level (ρ ≈ 0.87). Similarly, the number of sessions showed a negative correlation with post‑training divergence (ρ ≈ −0.47 for individuals and ≈ −0.997 for group averages): more sessions went hand‑in‑hand with better “outcomes,” even in toys with no brain activity at all.

Interestingly, the text of the discussion suggests that adolescents improved the most and elders the least, while the numerical table appears to show the largest pre‑post difference in the “elderly” group. This small internal inconsistency underlines the authors’ main point: when the participants are plush, any attempt at fine‑grained interpretation starts to feel slightly absurd.

Ultimately, the authors conclude that because stuffed animals lack brain electrical activity—and any ambient electrical noise in the environment should be relatively constant—systematic age effects and pre‑post improvements are difficult to reconcile with a straightforward reading of the divergence index as a measure of genuine brain self‑regulation.


Discussion

Beneath the humor of wiring up teddy bears to a neurofeedback system lies a serious methodological question: how do we know that a given neurofeedback metric reflects genuine brain activity rather than software behavior? This study uses satire as a stress test. By applying NeurOptimal to inert matter and still finding systematic improvements and age effects, the authors highlight the need for robust validation of black‑box systems.

One of the most striking features here is how “real” the data look. Pre‑post improvements, age‑related differences, and dose–response relationships (more sessions, better outcomes) are exactly the kinds of patterns clinicians love to see in their human clients. If the same patterns can be produced in stuffed animals simply by changing age fields in the software and repeating sessions, we are forced to ask: how much of what we see in clinical dashboards is about the brain, and how much is about the algorithm?

For people considering neurofeedback, this doesn’t mean that all systems are fake or that their genuine improvements are imaginary. What this paper satirizes is not neurofeedback as a whole, but specifically black‑box products like NeurOptimal that borrow the language of neurofeedback without offering the core elements that make it a scientific intervention: qEEG‑ or assessment‑guided targets, clearly defined EEG bands and sites, and transparent signal processing. Bodies and brains are exquisitely trainable; years of research using EEG‑based protocols, fMRI neurofeedback, and classic biofeedback (like heart rate variability training) show meaningful changes in symptoms, cognition, and physiology. What this teddy‑bear experiment calls out is the risk of over‑trusting any device that wraps complex internal processing in a simple, comforting summary score like “divergence,” especially when that score is opaque and proprietary.

For clinicians referring to neurofeedback services, the study is a gentle reminder to ask a few hard questions before sending patients along. Does the system provide access to raw EEG or physiologic data, or only to a black‑box index? Are there peer‑reviewed studies showing that its metrics track real brain changes, ideally using independent equipment? Is improvement defined only by the device’s own numbers, or by converging evidence from symptoms, behavior, and external measures (like qEEG, cognitive tests, or standardized questionnaires)?

For neurofeedback practitioners, the message is perhaps the most pointed—especially for those who have wondered whether plug‑and‑play systems like NeurOptimal belong in the same category as qEEG‑based, protocol‑driven EEG‑neurofeedback. Many of us rely on observable, interpretable EEG features: excessive frontocentral theta associated with inattention, low sensorimotor rhythm (SMR) at C3/C4 in hyperactivity, or asymmetric frontal alpha in mood disorders. Traditional EEG‑neurofeedback protocols—especially when informed by qEEG—build explicit, testable bridges between these patterns, the training targets, and the client’s experience. When a system obscures this link—replacing concrete frequencies and sites with a single global variable that always seems to go down over time—we lose the ability to critically check whether the feedback is actually doing what it claims.

The broader theme here echoes findings from sham‑controlled neurofeedback research: some portion of clinical benefit likely comes from non‑specific factors such as therapist attention, relaxation time, expectancy, and the general therapeutic container. The teddy‑bear study cleverly strips away all of these human variables and shows that the device alone can generate an appealing improvement story. That doesn’t negate genuine neuroplastic changes seen in rigorous trials of protocols like SMR training for ADHD or alpha up‑training for anxiety, but it does remind us that devices which cannot be independently validated deserve extra scrutiny.

An interpretive way to read this paper is as a critique of “neuro‑theatre”—the tendency to wrap traditional supportive care in high‑tech aesthetics and brain‑shaped language without proportionate evidence. NeurOptimal, as portrayed here, is a prime example: sophisticated branding, confident claims, and a tidy global metric, but no way to verify that the feedback is actually contingent on meaningful brain activity. When clients (or practitioners) are dazzled by color‑coded dashboards and impressive‑sounding indices, it becomes easier to conflate the appearance of precision with actual mechanistic insight. Satirical designs like this one play an important role in re‑grounding the field: if your neurofeedback system can successfully rehabilitate a plush octogenarian, it’s time to re‑examine what your metrics really mean.


Brendan’s perspective

I have to admit: the mental image of carefully marking electrode sites on a teddy bear’s head with Ten20 paste is going to stay with me for a while. It’s funny, yes—but it also captures something essential about where neurofeedback sometimes goes wrong. When we treat the device as a kind of magical brain optimizer, we risk forgetting the basics: signal quality, clear targets, physiological plausibility, and good old‑fashioned critical thinking.

In day‑to‑day clinical practice, my anchor is always the combination of three things: a coherent story of the person’s difficulties, observable EEG patterns (often from a qEEG or at least structured baseline recordings), and how the person adapts to the protocols we choose in response. For example, with a client who has impulsivity and motor restlessness, I might work with SMR (12–15 Hz) at sites like C4 or C3, inhibiting excessive theta (4–7 Hz) and high beta (22–30 Hz). For someone with anxious rumination and difficulty settling, I might favor increasing posterior alpha (around 9-12 Hz) at POz while gently inhibiting fast beta. These choices are not magic, they’re hypotheses grounded in decades of EEG research and updated continuously by observing how the person responds.

Now imagine replacing that entire process with a single number labeled “divergence” or “stability,” which reliably improves with every session regardless of what is actually on the scalp, and, as this study shows, even when there’s no brain under the electrodes at all. That’s essentially what this teddy‑bear study is poking at with NeurOptimal. That’s essentially what this teddy‑bear study is poking at. If stuffed animals show the same graceful downward curves in divergence that our human clients do, we should be very hesitant to use that number alone as proof of brain change.

This is where protocol individualisation becomes crucial. Two clients with the same diagnosis—say, ADHD—might have very different EEG presentations. One may show elevated frontal theta and low beta; another may have relatively normal theta/beta ratios but pronounced emotional reactivity with paroxysmal high‑beta bursts in right frontal regions. In the first case, a classic theta‑down, beta‑up protocol at Fz or Cz might be appropriate; in the second, I might work more with SMR at C4 and high‑beta inhibition at F4, plus complementary approaches like heart‑rate variability (HRV) training and emotion‑focused psychotherapy. A one‑size‑fits‑all global metric can’t tell those stories apart.

The satire also highlights how important it is to look at the raw signal. With any EEG system I’m willing to use clinically, I want to see the actual waveform, check impedance, and watch how artifacts—eye blinks, muscle tension, jaw clenching—show up in the recording. If all I ever see is a dashboard indicator that quietly glides from red to green, I’m flying blind. In contrast, well‑designed neurofeedback platforms let you see the data you are training, confirm that it looks like plausible human EEG, and relate changes in that signal to the person’s subjective experience. (Note: every single one of my training screens shows the raw EEG signal, and I will not ever work with a system that doesn’t.)

Another lesson from the teddy bears is about research design. Good neurofeedback trials increasingly include sham or non‑contingent feedback conditions, where the feedback is decoupled from the participant’s own brain activity. When improvements occur in both real and sham conditions, we learn that non‑specific factors are doing some of the heavy lifting. When only the real‑feedback group shows sustained changes in both physiology and behavior, we gain confidence that we’re tapping genuine brain‑based learning. Running a sham condition on stuffed animals is a wonderfully irreverent way of asking: “What part of this effect actually requires a brain?”

In terms of complementary methods, I’m a big fan of combining genuine EEG‑based neurofeedback—grounded in qEEG or at least structured assessment, with clear band‑ and site‑specific targets—with other psychophysiological tools. HRV biofeedback can support autonomic regulation and interoceptive awareness, while EMG biofeedback helps clients recognize and release chronic muscle tension. In many cases, these tools are more transparent than some of the more opaque neurofeedback systems: you can feel your breathing slow, see your heart‑rate oscillations smooth out, and experience the shift in state. When we add protocol‑based neurofeedback on top, targeting, say, increased alpha at Pz for a client with trauma‑related hypervigilance, or beta‑down training at F3/Fz for someone with racing thoughts, we’re working with a layered, intelligible model rather than a single mystical index.

What about applying lessons from this paper to other populations? One obvious takeaway is that whenever we work with vulnerable clients—children, individuals with neurodevelopmental conditions, people with progressive neurological diseases—we owe them extra rigor. That means using systems and protocols with traceable mechanisms: if we claim that training SMR at Cz can stabilize sleep onset and reduce nocturnal awakenings, we should be able to show changes in SMR power, sleep diaries, and (ideally) objective sleep measures. A metric that improves equally in teddy bears and teenagers simply doesn’t meet that bar.

And yet, I don’t see this satire as anti‑neurofeedback. If anything, it’s pro‑neurofeedback done well, and sharply critical of products like NeurOptimal that, in my view, are closer to a packaged relaxation/multimedia experience than to neurofeedback in the scientific sense. It nudges us away from blind trust in technology and back toward the fundamentals of learning, physiology, and therapeutic relationship. It reminds us that the goal isn’t to make a number on a screen look prettier—it’s to help real nervous systems find more flexible, resilient patterns of functioning.

So the next time you see a device that promises to “optimize your brain effortlessly” with minimal assessment and zero explanation of what’s actually being trained, you might quietly ask: would this also work on a teddy bear? If the honest answer is “probably,” then that system may belong more in the toy aisle than in a clinic.


Conclusion

The stuffed‑animal NeurOptimal study is funny on the surface, but its implications are serious for how we talk about what does—and does not—count as real neurofeedback. By showing that an automated neurofeedback system can produce convincing patterns of improvement and age‑related effects in inert matter, the authors force us to question how we validate proprietary metrics and interpret glossy dashboards. The findings don’t negate the large body of evidence supporting qEEG‑guided, protocol‑based EEG‑neurofeedback and other biofeedback methods as tools for self‑regulation and symptom change, but they do highlight the difference between transparent, mechanism‑based protocols and opaque, black‑box “optimization” promises of systems like NeurOptimal.

For clients, this is an invitation to ask better questions. For clinicians and researchers, it’s a gentle push toward systems that let us see and understand the signals we are training—and to include proper controls, including sham conditions, where appropriate. And for the field as a whole, it’s a reminder that our work is ultimately about living, dynamic nervous systems, not just software outputs.

When neurofeedback is grounded in good science, thoughtful assessment, and genuine human connection, it doesn’t need to impress teddy bears to change lives.


References

Pérez‑Elvira, R., & Carpena‑Niño, M. G. (2011). Scoping study on OPERATION NeurOptimal of Zengar [Unpublished manuscript]. Program Mental Rehabilitation, Hospital Care Center LAGUNA. Full text available here.

NEWSLETTER

Get the next one by email.

One article a week. No sequences, no pitch — just the reading, with the clinical read-out attached.

FREE WEEKLY SESSION

Bring the question to the room.

Let’s talk NEURO is an open hour for practitioners. Nothing is recorded for publication and nothing is sold — it exists so people can say what they are actually unsure about.

WRITTEN BY

Brendan Parsons, M.Sc., Ph.D., BCN

Neuroscientist and BCIA-certified practitioner, AAPB board member and Education Committee chair, BFE Service Award 2025. He founded NeuroLogic and has practised clinical qEEG for more than twenty years. More about the team

More in this thread

The Landscape of Neurofeedback Methods

Neurofeedback is not a single method. This guide maps the main families of EEG neurofeedback — from frequency-band training to z-score and LORETA approaches — and what the evidence says about each.

Read

Real-Time Source Imaging with hdEEG

RT-NET combines individualised head modelling, adaptive artefact attenuation, and source localisation to estimate source-space hdEEG activity online rather than only offline.

Read