---
license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
base_model: Qwen/Qwen3.8-27B
tags:
- qwen3_5
- agent
- deep-research
- reasoning
- tool-use
- long-context
- self-improvement
- safetensors
---
AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
## Introduction
AREX-2 is a 27B-parameter long-horizon agent model from the Beijing Academy of
Artificial Intelligence (BAAI). It learns to improve a solution over multiple
test-time rounds: propose, measure, reflect, and revise.
AREX-2 is trained on machine-learning and algorithmic-programming tasks with
verifiable feedback, together with the existing AREX deep-research data. The
learned self-improvement behavior transfers to deep research without adding new
search trajectories.
- **Architecture:** Dense Qwen3.8-compatible multimodal model
- **Parameters:** 27B
- **Context length:** 262,144 tokens
## Key features
- **Long-horizon self-improvement:** turns extra test-time rounds into useful
solution refinement.
- **Feedback-driven reflection:** reads scores, logs, errors, and timings to
decide what to change next.
- **Cross-domain performance:** training on coding and machine-learning tasks
also improves the model's deep-research performance.
- **Long-horizon reasoning:** sustains productive iteration as the task budget
grows.
## Evaluation
AREX-2 is evaluated on algorithmic programming, machine-learning engineering,
deep research, and general agentic reasoning. Results follow the protocols
reported in the AREX-2 paper.
### Coding and machine-learning engineering
| Model |
Params |
Frontier-CS |
MLE-Lite |
| Closed-weight models |
| GPT-5.6 Sol | - | 76.4 | 72.7 |
| Claude Opus 4.8 | - | 74.5 | 63.6 |
| GPT-5.5 | - | 72.1 | 68.2 |
| Gemini-3.1-Pro | - | 68.9 | - |
| Qwen3.7-Max | - | 61.9 | - |
| Open-weight models |
| Kimi-K3 | 2.8T | - | 72.7 |
| Naive-N0.5-Flash | 309B | - | 73.7 |
| DeepSeek-V4-Pro | 1.6T | 44.7 | 54.5 |
| DeepSeek-V4-Flash | 284B | 39.1 | 51.5 |
| Kimi-K2.7-Code | 1T | 54.7 | - |
| GLM-5.3-Flash | 320B | 50.4 | - |
| Kimi-K2.6 | 1T | 46.9 | 66.7 |
| Frontis-MA1-35B | 35B | - | 71.2 |
| BigBang-V1 | 35B | - | 59.1 |
| Qwen3.6-35B-A3B | 35B | 23.4 | 39.4 |
| AREX-2 | 27B | 70.7 | 81.8 |
Frontier-CS is the 188-task Agent Track. MLE-Lite reports Any Medal
averaged over three seeds.
### General agentic reasoning and deep research
| Model |
Params |
BrowseComp |
HLE |
GAIA |
DeepSearchQA |
| Frontier models |
| GPT-5.6 Sol | - | 90.4 | 58.0* | - | - |
| GPT-5.6 Terra | - | 87.5 | - | - | - |
| GPT-5.6 Luna | - | 83.3 | - | - | - |
| Kimi-K3 | 2.8T | 91.2 | 56.0* | - | 95.0 |
| Claude Fable 5 | - | 88.0 | 64.5* | - | 94.2 |
| Claude Opus 4.8 | - | 84.3 | 57.9* | - | 93.1 |
| GPT-5.5 | - | 84.4 | 52.2* | 87.4 | - |
| Gemini-3.1-Pro | - | 85.9 | 51.4* | 80.6 | 93.3 |
| Large models (>40B) |
| GLM-5 | 744B | 75.9 | 50.4 | 70.0 | - |
| Kimi-K2.6 | 1T | 83.2 | 54.0* | 80.6 | 92.5 |
| GLM-5.3-Flash | 320B | - | 55.3* | 78.8 | - |
| DeepSeek-V4-Flash | 284B | 73.2 | 45.1 | 57.5 | 90.6 |
| DeepSeek-V4-Pro | 1.6T | 83.4 | 48.2 | 71.1 | 88.7 |
| MiroThinker-1.7 | 235B | 74.0 | 42.9 | 82.7 | 72.1 |
| XYZ-Aquila-pro | 397B | 84.8 | 53.3 | - | 92.5 |
| Iris-pro | 397B | 88.6 | 56.4 | - | 92.9 |
| AREX-Base | 122B | 82.5 | 52.4 | 85.4 | 89.9 |
| Small models (≤40B) |
| Tongyi-DeepResearch-30B | 30B | 43.4 | 32.9 | 70.9 | - |
| Qwen3.5-35B | 35B | 61.0 | 47.4 | 80.0 | 68.5 |
| XYZ-Aquila-mini | 35B | 78.8 | 51.1 | 97.1 | 89.5 |
| BigBang-V1 | 35B | 76.5 | 50.3 | - | - |
| Quest-35B | 35B | 64.6 | 37.2 | 80.8 | - |
| Apodex-1.0-mini | 35B | 71.5 | 46.8 | - | 82.2 |
| Agents-A1 | 35B | 75.5 | 47.6 | 96.0 | - |
| MiroThinker-1.7-mini | 30B | 67.9 | 36.4 | 80.3 | 67.9 |
| Iris-mini | 35B | 82.2 | 52.3 | - | 86.9 |
| AREX-Turbo | 4B | 70.7 | 40.6 | 81.6 | 78.5 |
| AREX-2 | 27B | 84.0 | 52.6 | 92.2 | 93.8 |
HLE values marked * are from the full HLE set; unmarked values use the
text-only subset.
## Inference
Use a recent Transformers release with Qwen3.8 support.
```bash
pip install -U torch transformers accelerate
```
```python
import torch
from transformers import AutoModelForMultimodalLM, AutoProcessor
model_id = "BAAI/AREX-2"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForMultimodalLM.from_pretrained(
model_id, dtype=torch.bfloat16, device_map="auto"
)
messages = [{
"role": "user",
"content": "Propose a solution and explain how you would improve it over several rounds.",
}]
inputs = processor.apply_chat_template(
messages, add_generation_prompt=True, tokenize=True,
return_dict=True, return_tensors="pt",
).to(model.device)
with torch.inference_mode():
outputs = model.generate(**inputs, max_new_tokens=1024)
print(processor.decode(
outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True
))
```
## Intended use
AREX-2 is intended for research on long-horizon agents, iterative problem
solving, machine-learning engineering, algorithmic coding, and tool-augmented
deep research.
## License
AREX-2 is released under the [Apache License 2.0](LICENSE). Follow the terms
and notices for the Qwen base model and any downstream data or tools.
## Citation
```bibtex
@article{2026arex2,
title = {AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks},
author = {Qian, Hongjin and Li, Chaofan and Luo, Kun and Wei, Wenqing and Chen, Jianlyu and Lu, Shuqi and Hu, Yuyang and Xiao, Hongwang and Wang, Hui and Li, Chaozhuo and Ye, Qiwei and Dou, Zhicheng and Lian, Defu and Liu, Zheng},
journal = {arXiv preprint arXiv:2609.38288},
year = {2026},
url = {https://arxiv.org/abs/2609.38288}
}
```