Skip to content

AntResearchNLP/AlignXplorePlus

Folders and files

NameName
Last commit message
Last commit date

Latest commit

Β 

History

35 Commits
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

AlignXplore+

Text as a Universal Interface for Personalizing Large Language Models

arXiv πŸ€— HuggingFace πŸ€— HuggingFace πŸ€— HuggingFace πŸ€— HuggingFace

πŸ“– Introduction

AlignXplore+ is a framework for transferable personalization in large language models. Building upon AlignXplore, it represents a substantial evolution of the approach by advancing a central idea: natural language can serve as a universal, model- and task-agnostic interface for expressing fine-grained and multi-dimensional user preferences.

Model

✨ Key Features

  • General-Purpose : AlignXplore+ operates in a more realistic, real-world scenario, demonstrating that high-quality user preference summaries can be inferred from heterogeneous sources, including social networks, e-commerce platforms, and news streams.
  • Transferable : Preference summaries inferred by AlignXplore+ demonstrate strong transferability across both tasks (e.g., from response selection to news recommendation) and models (e.g., from GPT-OSS-20B to Qwen2.5-7B-Instruct).
  • Streaming & Robust : Inherited from AlignXplore, AlignXplore+ supports preference reasoning from streaming inputs and maintains stable performance under noisy or imperfect signals, ensuring reliable personalization in realistic, dynamic settings.

πŸ“Š Results

Main Results

The following is our main results across nine benchmarks. P-Soups are split into three preference dimensions: expertise, informativeness, and style. We compare models in three settings: direct sequence modeling, full-history preference inference, and streaming preference inference. All preference inference models (both full-history and streaming) use the Qwen3-8B-non-thinking as the downstream model. Within each setting, Bold and Underline mark the best and second-best results among models at the ~8B scale. Gray score highlights models that are outperformed by the best-performing ~8B model in the same column, including direct sequence models and larger preference inference models that do not show a performance advantage.

Model In-domain Out-of-domain Avg.
MIND Amazon AlignX MovieLens PersonaMem Info. Style Expertise HiCUPID
(Rec.) (Rec.) (R.S.) (Rec.) (R.S.) (R.S.) (R.S.) (R.S.) (R.G.)
Direct Full-history Sequence Models w/o Preference Inference
Qwen3-8Bnon-thinking 63.03 84.05 59.63 88.57 61.40 46.84 42.33 38.33 47.02 59.02
TALLRec 81.96 94.91 66.30 97.90 64.36 51.66 70.16 60.16 47.41 70.53
Full-history Preference Inference
DeepSeek-R1-671B 65.53 82.15 65.90 82.76 61.44 72.59 85.66 82.33 63.90 73.58
Qwen3-32Bthinking 67.63 85.69 64.93 75.43 57.36 73.25 88.00 83.66 63.44 73.26
GPT-OSS-20B 64.16 83.75 55.63 74.46 61.74 68.77 86.00 81.66 62.00 70.90
Qwen3-8Bthinking 66.10 84.68 62.73 75.13 54.36 75.08 87.50 83.50 60.05 72.12
DS-R1-Distill-Qwen-7B 61.20 82.82 54.03 70.30 49.28 56.14 65.83 66.00 60.01 62.84
AlignXplore 61.23 78.58 66.60 69.93 53.98 76.24 78.00 72.66 53.50 66.07
AlignXplore+ 71.36 86.39 75.03 75.80 58.08 78.07 86.33 82.50 62.42 75.10
Streaming Preference Inference
DeepSeek-R1-671B 64.30 80.54 64.06 83.63 58.96 66.61 85.00 79.00 60.32 72.93
Qwen3-32Bthinking 66.60 85.35 64.60 77.78 53.26 73.58 83.66 81.67 59.83 71.81
GPT-OSS-20B 64.93 84.55 56.86 73.66 54.82 69.93 83.00 77.50 59.93 69.46
Qwen3-8Bthinking 66.13 83.58 62.90 75.97 51.68 74.08 85.00 82.66 59.17 71.24
DS-R1-Distill-Qwen-7B 61.16 81.78 56.40 69.63 46.64 58.63 60.83 64.16 59.29 62.05
AlignXplore 60.66 79.01 69.90 67.96 48.42 74.41 74.83 69.16 50.34 66.07
AlignXplore+ 71.80 85.35 73.67 77.23 54.58 76.57 80.33 78.50 60.51 73.17

Cross-model Transfer Results

We evaluate if the generated summaries are useful for a variety of downstream models, not just the one used for training. Concretely, we first use the baselines and our AlignXplore+ model to generate user preference summaries, and then feed these summaries into Qwen2.5-7B-Instruct and GPT-OSS-20B to perform downstream tasks.

Qwen2.5-7B-Instruct
Model MIND Amazon AlignX MovieLens Info. Style Expertise PersonaMem AVG
Direct Full-history Sequence Models w/o Preference Inference
Qwen-2.5-7B-Instruct 52.60 67.75 52.96 91.30 50.83 40.16 51.33 40.76 55.96
Full-history Preference Inference
DeepSeek-R1-671B 63.80 79.93 64.10 79.33 66.44 76.16 76.50 64.96 71.40
Qwen3-32Bthinking 65.90 85.85 62.83 74.30 67.60 69.67 77.33 61.36 70.61
GPT-OSS-20B 61.73 85.05 55.86 75.23 68.27 68.33 75.33 65.74 69.44
Qwen3-8Bthinking 63.36 85.12 60.46 74.20 64.11 73.83 72.83 60.58 69.31
DS-R1-Distill-Qwen-7B 58.40 82.35 53.50 69.53 49.66 51.16 57.66 57.08 59.92
AlignXplore 57.53 81.12 65.20 69.83 67.94 63.00 68.16 53.60 65.80
AlignXplore+ 68.13 86.25 73.90 73.96 68.77 73.66 74.33 60.24 72.41
Streaming Preference Inference
DeepSeek-R1-671B 63.50 81.08 64.50 80.33 69.10 75.33 77.00 62.16 71.62
Qwen3-32Bthinking 65.16 85.35 63.56 74.56 68.27 70.67 73.50 57.72 69.85
GPT-OSS-20B 63.10 84.22 57.70 72.43 69.76 64.16 73.00 59.28 67.96
Qwen3-8Bthinking 64.13 84.12 61.10 72.14 67.94 75.83 76.09 56.46 69.73
DS-R1-Distill-Qwen-7B 58.10 82.55 55.56 69.86 51.99 48.66 59.16 54.22 60.01
AlignXplore 57.96 80.38 69.90 67.20 65.44 58.83 63.83 49.16 64.09
AlignXplore+ 67.73 85.05 74.00 75.56 71.92 71.33 74.00 55.32 71.86
GPT-OSS-20B
Model MIND Amazon AlignX MovieLens Info. Style Expertise PersonaMem AVG
Direct Full-history Sequence Models w/o Preference Inference
GPT-OSS-20B 63.86 88.99 68.50 85.26 79.40 83.66 84.50 21.92 72.01
Full-history Preference Inference
DeepSeek-R1-671B 60.01 80.00 68.60 77.33 76.24 90.00 85.00 63.86 75.13
Qwen3-32Bthinking 69.66 87.52 58.13 76.83 69.10 84.16 82.83 62.18 73.80
GPT-OSS-20B 66.43 86.29 47.40 75.86 70.93 84.16 85.16 59.52 71.97
Qwen3-8Bthinking 66.96 86.19 52.23 75.36 69.93 84.66 79.50 60.36 71.90
DS-R1-Distill-Qwen-7B 61.26 84.25 47.53 69.56 56.81 59.50 67.83 58.32 63.13
AlignXplore 61.60 82.35 64.73 69.13 73.75 77.16 73.50 59.48 70.21
AlignXplore+ 73.36 87.35 69.90 77.83 74.25 83.16 79.83 65.56 76.41
Streaming Preference Inference
DeepSeek-R1-671B 60.03 77.91 68.80 79.43 77.77 87.33 84.16 60.98 74.55
Qwen3-32Bthinking 69.43 86.89 56.26 77.10 67.27 81.50 80.83 58.88 72.27
GPT-OSS-20B 66.63 85.49 51.53 73.93 72.92 83.83 83.50 63.13 72.62
Qwen3-8Bthinking 67.63 85.09 52.90 75.13 69.10 82.00 78.16 56.48 70.81
DS-R1-Distill-Qwen-7B 60.43 84.95 50.63 69.33 53.15 54.50 66.83 52.12 61.49
AlignXplore 62.56 82.82 68.06 68.60 70.76 74.16 69.16 55.36 68.94
AlignXplore+ 73.20 87.75 68.66 78.93 74.41 81.00 75.00 61.22 75.02

Cross-task Transfer Results

We test a critical scenario: can a preference summary, derived from a user's behavior in one domain (e.g., recommendation), be successfully applied to guide personalization for the {same user} in a completely different domain (e.g., dialogue)? We evaluate this under full-history preference inference setting.

Model R.S. β†’ Rec. Rec. β†’ R.S. AVG
TALLRec 49.90 49.80 49.85
Qwen3-8Bthinking 57.90 50.40 54.15
AlignXplore+ 74.90 51.10 63.00

πŸš€ Quick Start

Requirments

You can install the required packages by running:

pip install -r requirements.txt

SFT Training

You should first download the dataset from here, then generate the tokenized dataset by running the following script.

cd sft

python prepare_dataset.py

You can modify line 21 and line 34 to set the path to your own model and tokenized dataset.

> sft/sft.py

21  model_name_or_path = "Qwen/Qwen3-8B" 

34  dataset = load_from_disk("tokenized_dataset")

Then set the node address and other distrubuted training parameters in the following script.

cd sft

bash sft.sh

RL Training

You should first download the RL dataset from here and run the following script to generate verl-format dataset.

cd verl

# set line 49 data_source_train = "path to jsonl data" to your downloaded dataset path.

python example/data_preprocess/upi_streaming_dataset.py

Then you can set your own path and run the RL training in the following script.

cd verl

bash examples/grpo_trainer/run_streaming_8p.sh

Inference

cd eval

# You should modify line #21 to the path of your model.
bash gen_preference.sh
bash straming_gen_preference.sh
# You should modify line #21 to the model name of your model.
bash evaluation.sh

βœ’οΈ Citation

@misc{liu2026textuniversalinterfacetransferable,
      title={Text as a Universal Interface for Transferable Personalization}, 
      author={Yuting Liu and Jian Guan and Jia-Nan Li and Wei Wu and Jiang-Ming Yang and Jianzhe Zhao and Guibing Guo},
      year={2026},
      eprint={2601.04963},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2601.04963}, 
}

About

Text as a Universal User Representation for Transferable Personalization

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages