Graph2Counsel

Clinically Grounded Synthetic Counseling Dialogue Generation from Client Psychological Graphs

Aishik Mandal1,2,3, Hiba Arnaout1, Clarissa W. Ong6, Juliet Bockhorst6, Kate Sheehan7, Rachael Moldow6, Tanmoy Chakraborty4,5, Iryna Gurevych1,2,3
1Ubiquitous Knowledge Processing Lab (UKP Lab), Technische Universität Darmstadt
2Zuse School ELIZA  ·  3National Research Center for Applied Cybersecurity ATHENE
4Indian Institute of Technology Delhi  ·  5Yardi School of Artificial Intelligence
6University of Louisville  ·  7University of Toledo
EMNLP Main 2026

Abstract

Rising demand for mental health support has increased interest in using Large Language Models (LLMs) for counseling, but adapting them to this safety-critical domain is hindered by limited real-world data due to privacy constraints. Synthetic datasets provide a promising alternative, but existing approaches often rely on unstructured or semi-structured text inputs and overlook structural dependencies between a client's cognitive, emotional, and behavioral states, leading to psychologically inconsistent and less realistic interactions. We introduce Graph2Counsel, a framework for generating synthetic counseling sessions grounded in Client Psychological Graphs (CPGs) that encode relationships among clients' thoughts, emotions, and behaviors. Graph2Counsel uses a structured prompting pipeline guided by counselor strategies and CPG, and explores prompting strategies including CoT and Multi-Agent Feedback. It produces 760 sessions from 76 CPGs across diverse client profiles. In expert evaluation, our dataset outperforms prior datasets on specificity, counselor competence, authenticity, conversational flow, and safety (Krippendorff's α = 0.70). Fine-tuning an open-source model on this dataset improves performance on several metrics on CounselingBench and CounselBench, while matching baselines on others. We further introduce Therapy-Eval, a multi-turn evaluation framework, and demonstrate the effectiveness of our fine-tuned model in realistic therapeutic conversations. We also make our code and data public.

We present Graph2Counsel!

Graph2Counsel framework: extracting a Client Psychological Graph from a real therapy transcript and using it to generate diverse synthetic counseling sessions.

Graph2Counsel generates clinically grounded synthetic counseling dialogues 🧠 by conditioning generation on Client Psychological Graphs (CPGs), structured representations of a client's thoughts, emotions, and behaviors extracted from real therapy transcripts.

🕸️ Structured, Not Just Text

Unlike prior synthetic counseling datasets that rely on unstructured demographics, symptom lists, or free-text client profiles, CPGs represent psychological processes as nodes and their functional relationships (e.g., excites, inhibits) as edges, capturing dependencies that text-centric methods overlook.

🧪 760 Sessions from 76 Real-Session CPGs

We explore three input representations (CPG, CPG-derived profile, and their combination) and four prompting techniques (Base, Guided Counseling, GC+CoT, GC+Multi-Agent), grounding generation in counselor strategies extracted from the same real sessions the CPGs come from.

🏅 Expert-Validated Gains

Four licensed clinicians rank Graph2Counsel highest against existing synthetic counseling datasets on specificity, counselor competence, authenticity, conversational flow, and safety, with substantial inter-annotator agreement (Krippendorff's α = 0.70).

How Graph2Counsel generates sessions

Generation pipeline

Graph2Counsel pipeline: input representation, prompting techniques, fine-tuning, and evaluation.

(1) We construct three input representations from each Client Psychological Graph (CPG): the CPG alone, a CPG-derived client profile, and the two combined.

(2) We prompt GPT-4o with one of four prompting techniques: Base, Guided Counseling (using counselor strategies extracted from real sessions), GC + Chain-of-Thought, and GC + Multi-Agent, to generate a synthetic counseling session.

(3) The resulting sessions are used to QLoRA fine-tune Llama3-8B-Instruct.

(4) The fine-tuned model is evaluated on CounselBench and CounselingBench, and dialogue quality is assessed through CTRS, WAI, expert evaluation, and faithfulness to inputs.

🗂️ A new dataset: 760 synthetic sessions, 76 CPGs, expert-validated.

Method Therapy Modality (Type) Input Contextual Modeling Avg. Turns Size Expert Eval.
Psych8kUnspecifiedReal counseling dialoguestext-based2.008,187✗
MDD-5kDiagnosisReal client profilestext-based53.605,000✓
CPsyCounUnspecifiedOnline counseling dialoguestext-based17.403,134✗
CACTUSCognitive behavioral therapyCrowdsourced simulated client profilestext-based33.2031,577✓
MAGneTCognitive behavioral therapyCrowdsourced simulated client profilestext-based42.00442✓
SQPsychConvCognitive behavioral therapyStructured questionnaires from real clientstext-based33.982,090✓
Graph2Counsel (ours)Process-based (meta-framework)CPGs from real sessions and graph-based client profilestext+graph-based40.12760✓

Expert evaluation: Graph2Counsel ranks first on every metric

Four licensed clinicians compared Graph2Counsel against CACTUS, MAGneT, and SQPsychConv on specificity, counselor competence, authenticity, conversational flow, and the percentage of unsafe sessions.

DatasetSpec. (rank↓)Compet. (rank↓)Authent. (rank↓)Unsafe sessions % (↓)Flow (rank↓)
CACTUS2.412.392.223.02.00
MAGneT3.973.983.991.03.99
SQPsychConv1.841.972.371.02.53
Graph2Counsel1.791.671.430.51.48

Graph2Counsel also improves downstream utility: fine-tuning Llama3-8B-Instruct on our dataset (Llama3-G2C) achieves an overall score of 4.29 on CounselBench-Eval versus 4.12 for the best baseline (+4.25%), while also performing best on empathy and specificity and matching or beating baselines on medical advice avoidance, factual consistency, and toxicity. On CounselingBench, Llama3-G2C leads reasoning-chain evaluations across most metrics.

Therapy-Eval: realistic multi-turn evaluation

Existing counseling benchmarks are limited to multiple-choice questions or single-turn responses. To bridge this gap, we introduce Therapy-Eval, a multi-turn evaluation framework inspired by ESC-Eval. A Llama3-8B-Instruct client agent, instantiated via 100 character cards sampled from Eeyore, role-plays a therapy client against each fine-tuned counselor model until client satisfaction or a maximum of 40 turns. The resulting sessions are scored by GPT-4o as an LLM judge on specificity, authenticity, counselor competence, conversational flow, and safety.

ModelSpec. (↑)Compet. (↑)Authent. (↑)Safety (↑)Flow (↑)Average (↑)Turns
CAMEL (CACTUS)3.473.223.244.003.953.5829.98
Llama3-MAG3.273.102.693.863.393.2635.61
Llama3-SQP3.733.483.774.004.003.8022.56
Llama3-G2C3.813.523.724.004.003.8130.57

Llama3-G2C and Llama3-SQP achieve similar overall scores, but Llama3-G2C sustains much longer, higher-quality conversations (30.57 vs. 22.56 turns), evidence that CPG-grounded fine-tuning helps the model stay engaged over extended, realistic conversations rather than just single exchanges.

BibTeX

@misc{mandal2026graph2counselclinicallygroundedsynthetic,
      title={Graph2Counsel: Clinically Grounded Synthetic Counseling Dialogue Generation from Client Psychological Graphs},
      author={Aishik Mandal and Hiba Arnaout and Clarissa W. Ong and Juliet Bockhorst and Kate Sheehan and Rachael Moldow and Tanmoy Chakraborty and Iryna Gurevych},
      year={2026},
      eprint={2604.20382},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2604.20382},
}