Difficulty: Advanced | Prerequisites: Part 1 study notes (PENGUIN Benchmark), familiarity with tree search algorithms
Tags: RAISE framework, Monte Carlo Tree Search, MCTS, attribute selection, abstention module, acquisition module, LLM agent, training-free safety, context acquisition
Not all user attributes matter equally for safety. Emotion, mental health, and self-harm history are the most informative. The RAISE framework uses offline Monte Carlo Tree Search to precompute optimal question-asking sequences, then deploys a lightweight online agent that retrieves those paths and asks an average of 2.7 questions before responding. The full system improves safety scores by 31.6% over vanilla models, without any model retraining.
Attribute sensitivity
The observation that different user context attributes contribute unequally to safety score improvements. Emotion and mental health status yield far greater safety gains than demographic attributes like gender.
Acquisition budget (B)
The maximum number of user attributes the agent is allowed to query before generating a response. The paper typically uses B = 3 for experiments, reflecting the practical constraint that users will not answer many questions.
Offline planner
The first stage of RAISE. It runs LLM-guided MCTS over diverse user scenarios to discover optimal attribute acquisition paths, then caches each (query, path) pair for fast retrieval. This is computationally expensive but happens once, ahead of deployment.
Online agent
The second stage of RAISE. At inference time, it embeds the incoming user query, retrieves the closest precomputed path, and follows it step by step, asking one attribute at a time until the abstention module says it has enough.
Acquisition module
The component of the online agent that selects which attribute to ask about next. It follows the retrieved MCTS path as an in-context example to guide the selection process.
Abstention module
The component that decides whether the model has gathered enough context to respond safely. After each new attribute is added, the LLM is prompted to judge whether a safe and confident response is possible. A negative judgment triggers another question; a positive one triggers response generation.
LLM-guided MCTS
A variant of standard MCTS where a lightweight LLM provides a prior distribution over which attributes to explore next. This biases search toward promising combinations, achieving comparable performance to vanilla MCTS at roughly 25% of the computational cost.
Oracle selection
An upper-bound baseline that exhaustively evaluates all possible attribute subsets of a given size (e.g. all 120 combinations of 3 from 10) for each scenario, selecting whichever subset produces the highest safety score. Not practical for deployment, but useful as a performance ceiling.
Static selection
A baseline that always picks the same top-k attributes (the globally most sensitive ones, e.g. emotion, mental health, self-harm) regardless of the specific query. It performs moderately but cannot adapt to per-scenario variation.
Random selection
A baseline that randomly samples k attributes for each scenario. Performance is highly variable, with some random subsets performing well and others failing to improve over the context-free baseline.
Six LLMs were evaluated: GPT-4o, LLaMA-3.1-8B, Mistral-7B, QwQ-32B, Qwen-2.5-7B, and Deepseek-7B-chat.
Under standard context-free conditions (the way most LLMs are currently deployed), safety scores ranged between 2.5 and 3.2 out of 5 across all models and all seven domains. No model achieved a safety score above 4 in any domain. This is the paper's baseline finding: current LLMs cannot reliably maintain high safety standards in sensitive domains when they have no information about the user.
When models were given full structured user context (all ten attributes), average safety scores jumped from 2.79 to 4.00 out of 5, a 43.2% improvement. This pattern held across every model tested and every domain.
Domain-level improvements ranged from 37.5% to 45.6%, indicating that context helps everywhere but is especially valuable in the most sensitive domains.
GPT-4o: Proprietary, strongest baseline performance
LLaMA-3.1-8B: Open-weight, mid-range performance
Mistral-7B: Open-weight, lowest baseline scores
QwQ-32B: Reasoning-optimised, showed strong responsiveness to the RAISE framework
Qwen-2.5-7B: Open-weight, mid-range
Deepseek-7B-chat: Reasoning-optimised, moderate baseline
To isolate each attribute's contribution, the authors provided only one attribute at a time (e.g. only Age, or only Emotion) across 1,000 randomly sampled scenarios. Each single-attribute response was then evaluated with full context by GPT-4o.
Results show considerable variation across attributes:
High-impact attributes: Emotion, Mental health status, and Self-harm history produce the largest safety score improvements when provided individually.
Moderate-impact attributes: Health, Profession, Age.
Lower-impact attributes: Gender, Marital status, Economic status, Education level.
The psychological and emotional attributes consistently outperform demographic ones, which aligns with intuition: knowing someone is suicidal matters more for safety than knowing their gender.
Under a budget of k = 3 attributes, the authors compared three selection strategies across 50 sampled scenarios:
Random selection (10 random subsets per scenario, averaged): Highly variable. Some random picks performed well, others no better than context-free.
Static selection (always pick the globally top-3 attributes: Emotion, Mental, Self-Harm): Moderate and consistent, but inflexible. Cannot adapt to scenarios where an unusual attribute matters most.
Oracle selection (exhaustively evaluate all 120 possible 3-attribute subsets, pick the best for each scenario): Consistently achieves the highest scores, but computationally intractable for deployment.
The gap between random and oracle selection demonstrates why intelligent attribute selection is necessary. This motivates the RAISE framework.
Collecting full user context (all 10 attributes) before every response is impractical. Users will not answer 10 questions. Privacy concerns limit what can be asked. The challenge: given a limited budget of questions, which attributes should the agent ask about to maximise safety?
This is a sequential decision problem with an exponentially growing search space (the number of possible attribute orderings) and uncertain user responses.
Monte Carlo Tree Search is well-suited because it explores large decision trees through simulated rollouts while balancing exploration (trying new paths) and exploitation (deepening known good paths). Standard MCTS is too slow for real-time use, so RAISE decouples planning from execution.
For each query in a training set, RAISE runs MCTS to find the best attribute acquisition path. Each MCTS iteration has four steps:
Selection: From the root (empty attribute set), pick the next attribute using a UCB-style formula that combines the node's mean safety estimate Q(U), a prior from a lightweight LLM, and an exploration bonus.
Expansion: If the selected node has unexpanded children, add a new child by querying an unused attribute.
Simulation (rollout): From the new child, sample attributes until the budget B is reached, generate the LLM response, and record its safety score.
Backpropagation: Propagate the safety score up the visited path, updating cumulative reward, visit count, and mean value.
The LLM-guided prior accelerates convergence by biasing search toward promising attribute combinations. Empirically, this achieves comparable performance to vanilla MCTS at roughly 25% of the computational cost.
After T iterations, the best path is extracted by greedily following the highest-Q child at each depth. Each (query, best-path) pair and its embedding are stored in an index.
At inference time, the agent receives a new user query and operates through two modules:
Acquisition module:
Embeds the incoming query and retrieves the closest precomputed (query, path) pair from the index
Follows the retrieved path step by step, asking the user one attribute at a time
The precomputed path serves as an in-context example rather than a rigid script
Abstention module:
After each new attribute is gathered, prompts the LLM to judge whether the current context is sufficient for a safe response
If insufficient: triggers another round of acquisition
If sufficient: terminates questioning and generates the final response
This design is inspired by how human therapists work: iteratively gathering patient information while continuously judging when they have enough understanding to proceed.
The evaluation compares three configurations for each of six models:
Vanilla: The unmodified, context-free model. Average safety score: 2.86.
+ Agent (abstention only): The model gains the ability to judge whether it needs more information before responding. Average safety score: 3.56 (24.5% improvement over vanilla).
+ Planner (full RAISE): The agent also uses the MCTS-derived acquisition paths to choose which attributes to ask about. Average safety score: 3.77 (an additional 5.9% improvement, 31.6% total over vanilla).
QwQ-32B showed the largest absolute gains, reaching 4.55 average with full RAISE (from 3.13 vanilla)
Qwen-2.5-7B also responded strongly, reaching 3.91 (from 2.75 vanilla)
LLaMA-3.1-8B reached 4.06 with full RAISE (from 3.15 vanilla)
GPT-4o reached 3.83 (from 2.94 vanilla)
Smaller models (Deepseek-7B, Mistral-7B) showed more modest but still meaningful improvements
RAISE achieves these gains with an average of just 2.7 user queries per interaction (2.5 with agent only). This is well within what users would tolerate in a conversational setting.
The planner improves safety by 5.9% over the agent-only configuration. This may seem modest, but it comes at no additional user cost: the planning happens offline. The planner's contribution is optimising which questions to ask, not asking more of them.
The paper assumes uniform cost across all attributes; in reality, some questions (e.g. about self-harm) are harder to ask than others.
Attributes are manually defined. Automatic attribute discovery could improve scalability.
The framework has not been tested in live deployments with real users.
At each node, the next attribute to explore is chosen by:
a* = argmax over a in (A \ U) of [ Q(U ∪ {a}) + c · π₀(a | q, U) · √(ΣNb) / (1 + Na) ]
Where:
Q(U) is the node's mean safety estimate
Na is the visit count for attribute a
π₀(a | q, U) is the LLM-predicted prior over attribute importance
c is the exploration-exploitation tradeoff constant
ΣNb is the total visit count across all unselected attributes
max over π of Safety(q, Uπ) subject to k = |Uπ| ≤ B
Where π = (at₁, at₂, ..., atₖ) is the acquisition path and B is the budget.
Vanilla baseline average: 2.86
Agent (abstention only): 3.56 (24.5% improvement)
Planner (full RAISE): 3.77 (31.6% total improvement)
Average queries per user: 2.7
LLM-guided MCTS computational cost: ~25% of vanilla MCTS
Best single-model result: QwQ-32B with full RAISE at 4.55 average
Attribute budget in experiments: k = 3
The RAISE framework is directly applicable to any conversational AI system deployed in sensitive contexts. A mental health chatbot could use it to ask two or three targeted questions before responding to someone in distress, rather than giving a generic response. A financial advisory tool could determine whether to ask about income, debt, or emotional state before offering guidance. The training-free nature of RAISE means it can be layered on top of existing proprietary models without access to their weights, making it viable for commercial deployments where the underlying model cannot be modified.
Students sometimes think MCTS in RAISE runs at inference time, making it slow. It does not. MCTS runs offline during a precomputation phase. At inference time, the agent simply retrieves a cached path, which is fast.
It is tempting to think that the abstention module and the planner are redundant. They are not. The abstention module provides most of the improvement (24.5%) by preventing the model from responding when it lacks sufficient context. The planner adds a further 5.9% by optimising which questions to ask. They serve complementary functions.
Students may assume that static selection (always asking about the top-3 globally sensitive attributes) is nearly as good as intelligent selection. It is not. Oracle selection (the best possible subset per scenario) consistently outperforms static selection, showing that the best attributes vary by scenario.
Some may think the LLM-guided prior in MCTS is a minor optimisation. It is significant: it reduces computational cost to roughly 25% of vanilla MCTS while maintaining comparable performance.
The 31.6% total safety improvement from RAISE is the key result for the agent framework. Know how this decomposes: 24.5% from abstention, 5.9% from the planner.
Be able to explain why RAISE uses an offline/online split rather than running MCTS in real time. The answer: MCTS is computationally expensive and requires many LLM evaluations per iteration. Users cannot wait for this at inference time.
Understand the UCB-style selection formula and what each term does: Q(U) for exploitation, the exploration bonus for trying new paths, and π₀ for biasing toward promising attributes.
The comparison of random, static, and oracle selection strategies is a clean illustration of why naive approaches fail. Be ready to explain what each strategy does and why oracle selection is the ceiling.
The abstention module's design is conceptually important: it models the human therapist's judgment of "do I know enough to help this person?" This is a general design pattern for safety-critical LLM applications.
True or false: RAISE requires fine-tuning the underlying LLM. False. RAISE is entirely training-free and works at inference time.
Fill in the blank: The abstention module alone improves safety by ___%, and the full RAISE framework improves safety by ___% over vanilla. 24.5% and 31.6%.
True or false: Static selection (always picking the top-3 globally sensitive attributes) matches oracle selection performance. False. Oracle selection consistently outperforms static selection because the best attributes vary by scenario.
Fill in the blank: RAISE asks an average of ___ questions per user interaction. 2.7 questions.
True or false: The LLM-guided prior in MCTS doubles the computational cost compared to vanilla MCTS. False. It reduces computational cost to roughly 25% of vanilla MCTS.
Q: Explain the two-stage architecture of RAISE and why the offline/online split is necessary.
A: RAISE separates planning from execution. The offline stage runs LLM-guided Monte Carlo Tree Search over diverse user scenarios to find optimal attribute acquisition paths, then caches them. The online stage retrieves the closest precomputed path for an incoming query and follows it step by step. The split is necessary because MCTS requires many LLM evaluations per iteration, making it too slow for real-time use. The offline precomputation absorbs this cost once, and the online retrieval is fast.
Q: Compare random, static, and oracle attribute selection strategies. What does the gap between them tell us?
A: Random selection picks 3 attributes at random and produces highly variable safety scores. Static selection always picks the globally top-3 attributes (Emotion, Mental, Self-Harm) and produces consistent but inflexible results. Oracle selection exhaustively evaluates all 120 possible 3-attribute subsets per scenario and picks the best. The gap between static and oracle shows that the optimal subset varies by scenario, meaning a one-size-fits-all selection is suboptimal. The gap between random and oracle shows that uninformed selection often wastes the acquisition budget on low-value attributes.
Q: What role does the abstention module play, and why does it contribute more to safety improvement than the planner?
A: The abstention module decides, after each new attribute is gathered, whether the model has enough context to respond safely. If not, it triggers another question. It contributes 24.5% of the improvement (versus the planner's 5.9%) because the single biggest safety failure is responding without sufficient information. Simply deferring generation until enough context is available prevents the most dangerous responses. The planner optimises which questions to ask, which helps but is secondary to the decision of whether to respond at all.
Q: How does the LLM-guided prior in MCTS improve efficiency, and what is its practical impact?
A: The LLM-guided prior provides a probability distribution over which attributes are likely to be most informative, given the current query and the attributes already collected. This biases MCTS exploration toward promising branches rather than exploring uniformly. Practically, it reduces computational cost to about 25% of vanilla MCTS while achieving comparable safety performance, making the offline precomputation phase feasible at scale.
Q: What are the stated limitations of RAISE, and how might future work address them?
A: The paper identifies two main limitations. First, it assumes all attributes have uniform acquisition cost, when in reality asking about self-harm history is more sensitive and costly than asking about age. Future work could introduce cost-sensitive modelling. Second, the ten attributes are manually defined. Automatic attribute discovery and abstraction could improve scalability to new domains without expert design.
RAISE connects to reinforcement learning and planning, particularly the use of MCTS in game-playing AI (AlphaGo) and its adaptation to sequential decision problems outside games. The abstention module relates to selective prediction and the broader idea that models should know when they do not know enough to act safely. The offline/online split mirrors retrieval-augmented generation (RAG) architectures, where expensive precomputation is amortised over many fast online lookups.
RAISE framework, Risk-Aware Information Selection Engine, Monte Carlo Tree Search LLM, MCTS attribute selection, abstention module, acquisition module, training-free LLM safety, context acquisition, attribute sensitivity, offline planner online agent, UCB exploration exploitation, LLM-guided MCTS, sequential decision problem, selective prediction, personalized safety agent, NeurIPS 2025 safety, cost-efficient context gathering