AlignXplore+ is a framework for transferable personalization in large language models. Building upon AlignXplore, it represents a substantial evolution of the approach by advancing a central idea: natural language can serve as a universal, model- and task-agnostic interface for expressing fine-grained and multi-dimensional user preferences.
- General-Purpose : AlignXplore+ operates in a more realistic, real-world scenario, demonstrating that high-quality user preference summaries can be inferred from heterogeneous sources, including social networks, e-commerce platforms, and news streams.
- Transferable : Preference summaries inferred by AlignXplore+ demonstrate strong transferability across both tasks (e.g., from response selection to news recommendation) and models (e.g., from GPT-OSS-20B to Qwen2.5-7B-Instruct).
- Streaming & Robust : Inherited from AlignXplore, AlignXplore+ supports preference reasoning from streaming inputs and maintains stable performance under noisy or imperfect signals, ensuring reliable personalization in realistic, dynamic settings.
The following is our main results across nine benchmarks. P-Soups are split into three preference dimensions: expertise, informativeness, and style. We compare models in three settings: direct sequence modeling, full-history preference inference, and streaming preference inference. All preference inference models (both full-history and streaming) use the Qwen3-8B-non-thinking as the downstream model. Within each setting, Bold and Underline mark the best and second-best results among models at the ~8B scale. Gray score highlights models that are outperformed by the best-performing ~8B model in the same column, including direct sequence models and larger preference inference models that do not show a performance advantage.
| Model | In-domain | Out-of-domain | Avg. | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| MIND | Amazon | AlignX | MovieLens | PersonaMem | Info. | Style | Expertise | HiCUPID | ||
| (Rec.) | (Rec.) | (R.S.) | (Rec.) | (R.S.) | (R.S.) | (R.S.) | (R.S.) | (R.G.) | ||
| Direct Full-history Sequence Models w/o Preference Inference | ||||||||||
| Qwen3-8Bnon-thinking | 63.03 | 84.05 | 59.63 | 88.57 | 61.40 | 46.84 | 42.33 | 38.33 | 47.02 | 59.02 |
| TALLRec | 81.96 | 94.91 | 66.30 | 97.90 | 64.36 | 51.66 | 70.16 | 60.16 | 47.41 | 70.53 |
| Full-history Preference Inference | ||||||||||
| DeepSeek-R1-671B | 65.53 | 82.15 | 65.90 | 82.76 | 61.44 | 72.59 | 85.66 | 82.33 | 63.90 | 73.58 |
| Qwen3-32Bthinking | 67.63 | 85.69 | 64.93 | 75.43 | 57.36 | 73.25 | 88.00 | 83.66 | 63.44 | 73.26 |
| GPT-OSS-20B | 64.16 | 83.75 | 55.63 | 74.46 | 61.74 | 68.77 | 86.00 | 81.66 | 62.00 | 70.90 |
| Qwen3-8Bthinking | 66.10 | 84.68 | 62.73 | 75.13 | 54.36 | 75.08 | 87.50 | 83.50 | 60.05 | 72.12 |
| DS-R1-Distill-Qwen-7B | 61.20 | 82.82 | 54.03 | 70.30 | 49.28 | 56.14 | 65.83 | 66.00 | 60.01 | 62.84 |
| AlignXplore | 61.23 | 78.58 | 66.60 | 69.93 | 53.98 | 76.24 | 78.00 | 72.66 | 53.50 | 66.07 |
| AlignXplore+ | 71.36 | 86.39 | 75.03 | 75.80 | 58.08 | 78.07 | 86.33 | 82.50 | 62.42 | 75.10 |
| Streaming Preference Inference | ||||||||||
| DeepSeek-R1-671B | 64.30 | 80.54 | 64.06 | 83.63 | 58.96 | 66.61 | 85.00 | 79.00 | 60.32 | 72.93 |
| Qwen3-32Bthinking | 66.60 | 85.35 | 64.60 | 77.78 | 53.26 | 73.58 | 83.66 | 81.67 | 59.83 | 71.81 |
| GPT-OSS-20B | 64.93 | 84.55 | 56.86 | 73.66 | 54.82 | 69.93 | 83.00 | 77.50 | 59.93 | 69.46 |
| Qwen3-8Bthinking | 66.13 | 83.58 | 62.90 | 75.97 | 51.68 | 74.08 | 85.00 | 82.66 | 59.17 | 71.24 |
| DS-R1-Distill-Qwen-7B | 61.16 | 81.78 | 56.40 | 69.63 | 46.64 | 58.63 | 60.83 | 64.16 | 59.29 | 62.05 |
| AlignXplore | 60.66 | 79.01 | 69.90 | 67.96 | 48.42 | 74.41 | 74.83 | 69.16 | 50.34 | 66.07 |
| AlignXplore+ | 71.80 | 85.35 | 73.67 | 77.23 | 54.58 | 76.57 | 80.33 | 78.50 | 60.51 | 73.17 |
We evaluate if the generated summaries are useful for a variety of downstream models, not just the one used for training. Concretely, we first use the baselines and our AlignXplore+ model to generate user preference summaries, and then feed these summaries into Qwen2.5-7B-Instruct and GPT-OSS-20B to perform downstream tasks.
| Model | MIND | Amazon | AlignX | MovieLens | Info. | Style | Expertise | PersonaMem | AVG |
|---|---|---|---|---|---|---|---|---|---|
| Direct Full-history Sequence Models w/o Preference Inference | |||||||||
| Qwen-2.5-7B-Instruct | 52.60 | 67.75 | 52.96 | 91.30 | 50.83 | 40.16 | 51.33 | 40.76 | 55.96 |
| Full-history Preference Inference | |||||||||
| DeepSeek-R1-671B | 63.80 | 79.93 | 64.10 | 79.33 | 66.44 | 76.16 | 76.50 | 64.96 | 71.40 |
| Qwen3-32Bthinking | 65.90 | 85.85 | 62.83 | 74.30 | 67.60 | 69.67 | 77.33 | 61.36 | 70.61 |
| GPT-OSS-20B | 61.73 | 85.05 | 55.86 | 75.23 | 68.27 | 68.33 | 75.33 | 65.74 | 69.44 |
| Qwen3-8Bthinking | 63.36 | 85.12 | 60.46 | 74.20 | 64.11 | 73.83 | 72.83 | 60.58 | 69.31 |
| DS-R1-Distill-Qwen-7B | 58.40 | 82.35 | 53.50 | 69.53 | 49.66 | 51.16 | 57.66 | 57.08 | 59.92 |
| AlignXplore | 57.53 | 81.12 | 65.20 | 69.83 | 67.94 | 63.00 | 68.16 | 53.60 | 65.80 |
| AlignXplore+ | 68.13 | 86.25 | 73.90 | 73.96 | 68.77 | 73.66 | 74.33 | 60.24 | 72.41 |
| Streaming Preference Inference | |||||||||
| DeepSeek-R1-671B | 63.50 | 81.08 | 64.50 | 80.33 | 69.10 | 75.33 | 77.00 | 62.16 | 71.62 |
| Qwen3-32Bthinking | 65.16 | 85.35 | 63.56 | 74.56 | 68.27 | 70.67 | 73.50 | 57.72 | 69.85 |
| GPT-OSS-20B | 63.10 | 84.22 | 57.70 | 72.43 | 69.76 | 64.16 | 73.00 | 59.28 | 67.96 |
| Qwen3-8Bthinking | 64.13 | 84.12 | 61.10 | 72.14 | 67.94 | 75.83 | 76.09 | 56.46 | 69.73 |
| DS-R1-Distill-Qwen-7B | 58.10 | 82.55 | 55.56 | 69.86 | 51.99 | 48.66 | 59.16 | 54.22 | 60.01 |
| AlignXplore | 57.96 | 80.38 | 69.90 | 67.20 | 65.44 | 58.83 | 63.83 | 49.16 | 64.09 |
| AlignXplore+ | 67.73 | 85.05 | 74.00 | 75.56 | 71.92 | 71.33 | 74.00 | 55.32 | 71.86 |
| Model | MIND | Amazon | AlignX | MovieLens | Info. | Style | Expertise | PersonaMem | AVG |
|---|---|---|---|---|---|---|---|---|---|
| Direct Full-history Sequence Models w/o Preference Inference | |||||||||
| GPT-OSS-20B | 63.86 | 88.99 | 68.50 | 85.26 | 79.40 | 83.66 | 84.50 | 21.92 | 72.01 |
| Full-history Preference Inference | |||||||||
| DeepSeek-R1-671B | 60.01 | 80.00 | 68.60 | 77.33 | 76.24 | 90.00 | 85.00 | 63.86 | 75.13 |
| Qwen3-32Bthinking | 69.66 | 87.52 | 58.13 | 76.83 | 69.10 | 84.16 | 82.83 | 62.18 | 73.80 |
| GPT-OSS-20B | 66.43 | 86.29 | 47.40 | 75.86 | 70.93 | 84.16 | 85.16 | 59.52 | 71.97 |
| Qwen3-8Bthinking | 66.96 | 86.19 | 52.23 | 75.36 | 69.93 | 84.66 | 79.50 | 60.36 | 71.90 |
| DS-R1-Distill-Qwen-7B | 61.26 | 84.25 | 47.53 | 69.56 | 56.81 | 59.50 | 67.83 | 58.32 | 63.13 |
| AlignXplore | 61.60 | 82.35 | 64.73 | 69.13 | 73.75 | 77.16 | 73.50 | 59.48 | 70.21 |
| AlignXplore+ | 73.36 | 87.35 | 69.90 | 77.83 | 74.25 | 83.16 | 79.83 | 65.56 | 76.41 |
| Streaming Preference Inference | |||||||||
| DeepSeek-R1-671B | 60.03 | 77.91 | 68.80 | 79.43 | 77.77 | 87.33 | 84.16 | 60.98 | 74.55 |
| Qwen3-32Bthinking | 69.43 | 86.89 | 56.26 | 77.10 | 67.27 | 81.50 | 80.83 | 58.88 | 72.27 |
| GPT-OSS-20B | 66.63 | 85.49 | 51.53 | 73.93 | 72.92 | 83.83 | 83.50 | 63.13 | 72.62 |
| Qwen3-8Bthinking | 67.63 | 85.09 | 52.90 | 75.13 | 69.10 | 82.00 | 78.16 | 56.48 | 70.81 |
| DS-R1-Distill-Qwen-7B | 60.43 | 84.95 | 50.63 | 69.33 | 53.15 | 54.50 | 66.83 | 52.12 | 61.49 |
| AlignXplore | 62.56 | 82.82 | 68.06 | 68.60 | 70.76 | 74.16 | 69.16 | 55.36 | 68.94 |
| AlignXplore+ | 73.20 | 87.75 | 68.66 | 78.93 | 74.41 | 81.00 | 75.00 | 61.22 | 75.02 |
We test a critical scenario: can a preference summary, derived from a user's behavior in one domain (e.g., recommendation), be successfully applied to guide personalization for the {same user} in a completely different domain (e.g., dialogue)? We evaluate this under full-history preference inference setting.
| Model | R.S. β Rec. | Rec. β R.S. | AVG |
|---|---|---|---|
| TALLRec | 49.90 | 49.80 | 49.85 |
| Qwen3-8Bthinking | 57.90 | 50.40 | 54.15 |
| AlignXplore+ | 74.90 | 51.10 | 63.00 |
You can install the required packages by running:
pip install -r requirements.txtYou should first download the dataset from here, then generate the tokenized dataset by running the following script.
cd sft
python prepare_dataset.pyYou can modify line 21 and line 34 to set the path to your own model and tokenized dataset.
> sft/sft.py
21 model_name_or_path = "Qwen/Qwen3-8B"
34 dataset = load_from_disk("tokenized_dataset")Then set the node address and other distrubuted training parameters in the following script.
cd sft
bash sft.shYou should first download the RL dataset from here and run the following script to generate verl-format dataset.
cd verl
# set line 49 data_source_train = "path to jsonl data" to your downloaded dataset path.
python example/data_preprocess/upi_streaming_dataset.pyThen you can set your own path and run the RL training in the following script.
cd verl
bash examples/grpo_trainer/run_streaming_8p.shcd eval
# You should modify line #21 to the path of your model.
bash gen_preference.sh
bash straming_gen_preference.sh
# You should modify line #21 to the model name of your model.
bash evaluation.sh@misc{liu2026textuniversalinterfacetransferable,
title={Text as a Universal Interface for Transferable Personalization},
author={Yuting Liu and Jian Guan and Jia-Nan Li and Wei Wu and Jiang-Ming Yang and Jianzhe Zhao and Guibing Guo},
year={2026},
eprint={2601.04963},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2601.04963},
}