P(s | H)
WHERE
Locate the exact counselor utterance that contains an ethical violation.
AI supervision for counselor education
AI Supervisor Scaffolds Novice Growth in Counselor Education
We reposition AI from patient-facing counselor to educational supervisor—helping novices notice subtle ethical violations, understand their risks, and learn safer responses.
The idea
The most dangerous novice errors are often not obviously harmful. They can sound warm and supportive while quietly violating professional ethics.
We build an AI supervisor that does not replace novice counselors, but grows them. The supervisor locates an ethics-violating utterance, diagnoses it against APA principles, and explains not only what went wrong, but why it is risky and how to respond differently.
To overcome the lack of labeled clinical violations, a controllable AI novice intentionally enacts predefined mistake categories, making supervision labels a natural byproduct of generation. This yields EthicScaff, a 9,915-instance human-in-the-loop dataset. A Novice Growth Reward then optimizes the supervisor for whether a weaker novice model actually improves after reading its explanation.
Experiments show better downstream counseling behavior, sharper ethical detection, and significant self-efficacy gains across all eight assessed competencies in a study with novice counseling-psychology students.
Zone of proximal development
Expert therapeutic judgment is open-ended and contextual. Ethical boundaries are comparatively finite and teachable. The supervisor provides the missing scaffold between a novice who cannot recognize harm alone and a practitioner who has internalized professional constraints.
Novice commits subtle mistakes without recognizing the potential harm.
Scaffolded practice makes hidden violations visible and explainable.
Ethical growth turns explicit guidance into safer internalized practice.

Do No Harm supervision task
Given a counselor–client dialogue, the model produces a structured learning scaffold in three connected steps.
P(s | H)
Locate the exact counselor utterance that contains an ethical violation.
P(m | H, s)
Classify the violated principle and the novice-level mistake it reflects.
P(f | H, s, m)
Give pedagogical feedback that explains the risk and supports self-correction.
Method
At inference time the supervisor grows the novice. During training, novice growth becomes the signal that improves the supervisor.

Cross 106 client cases with 16 behavior patterns so subtle ethical violations become observable and labeled.
Validator-Guided Refinement and clinical expert review enforce progressivity, actionability, ethicality, and supportiveness.
Learn to localize violations, classify principles, and generate targeted explanatory feedback.
GRPO rewards explanations that improve a frozen novice before-versus-after reading the feedback—without leaking the answer.
Data construction
Clinical principles, controlled counselor–patient role-play, supervisor feedback, and expert review work together to produce high-quality dialogue–feedback pairs.
EthicScaff
Real counseling data rarely labels subtle ethical failures. EthicScaff makes these failures controllable, visible, and pedagogically useful while preserving realistic client variation.
Experiments
Evaluation spans downstream counselor behavior, objective ethical judgment, component ablations, expert assessment, and novice self-efficacy.
Qwen3-14B with the full framework, versus 38.67% for its base model and 73.12% for GPT-4o with RAG.
Qwen3-8B with the full framework achieves the best overall localization balance, including 63.03% Jaccard.
Compared with the unguided novice, gains are largest on MITI, WAI, and PSC—indicating better collaboration, alliance-building, and clinical appropriateness.
Professional feedback quality improves under automatic and expert judgment. Novice students report significant gains across every assessed counseling competency.



Responsible role
This work positions AI as an educational scaffold for low-stakes counselor training. It is not a patient-facing therapy system, a diagnostic tool, or a substitute for licensed clinical supervision. Its purpose is to help novices internalize professional constraints before working with vulnerable clients.
Citation
For the full methodology, experiments, and ethical considerations, see the arXiv manuscript.
Open PDF@article{xu2025first,
title = {First, Do No Harm: AI Supervisor Scaffolds
Novice Growth in Counselor Education},
author = {Xu, Chen and Lyu, Zhenyu and Lan, Tian and
Yi, Yang and Ji, Yu and Ji, Luyao and Shen, Jian and
Wang, Zhihua and Cui, Leyang and Zhang, Jieshuo and
Wan, Xiaohua and Dong, Qunxi and Yang, Minqiang and
Wang, Juan and Liu, Xiuling and Hu, Bin},
journal = {arXiv preprint arXiv:2508.09042},
year = {2025}
}