Aligning Compound AI Systems via System-level DPO

Wang, Xiangwen; Zhang, Yibo Jacky; Ding, Zhoujie; Tsai, Katherine; Wu, Haolun; Koyejo, Sanmi

Computer Science > Machine Learning

arXiv:2502.17721 (cs)

[Submitted on 24 Feb 2025 (v1), last revised 2 Dec 2025 (this version, v3)]

Title:Aligning Compound AI Systems via System-level DPO

Authors:Xiangwen Wang, Yibo Jacky Zhang, Zhoujie Ding, Katherine Tsai, Haolun Wu, Sanmi Koyejo

View PDF HTML (experimental)

Abstract:Compound AI systems, comprising multiple interacting components such as LLMs, foundation models, and external tools, have demonstrated remarkable improvements compared to single models in various tasks. To ensure their effective deployment in real-world applications, aligning these systems with human preferences is crucial. However, aligning the compound system via policy optimization, unlike the alignment of a single model, is challenging for two main reasons: (i) non-differentiable interactions between components make end-to-end gradient-based optimization method inapplicable, and (ii) system-level preferences cannot be directly transformed into component-level preferences. To address these challenges, we first formulate compound AI systems as Directed Acyclic Graphs (DAGs), explicitly modeling both component interactions and the associated data flows. Building on this formulation, we introduce $\textbf{SysDPO}$, a framework that extends Direct Preference Optimization (DPO) to enable joint system-level alignment. We propose two variants, SysDPO-Direct and SysDPO-Sampling, tailored for scenarios depending on whether we construct a system-specific preference dataset. We empirically demonstrate the effectiveness of our approach across two applications: the joint alignment of a language model and a diffusion model, and the joint alignment of an LLM collaboration system.

Comments:	NeurIPS 2025
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
Cite as:	arXiv:2502.17721 [cs.LG]
	(or arXiv:2502.17721v3 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2502.17721

Submission history

From: Yibo Zhang [view email]
[v1] Mon, 24 Feb 2025 23:25:13 UTC (18,673 KB)
[v2] Tue, 3 Jun 2025 20:03:48 UTC (19,113 KB)
[v3] Tue, 2 Dec 2025 05:21:20 UTC (19,120 KB)

Computer Science > Machine Learning

Title:Aligning Compound AI Systems via System-level DPO

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Aligning Compound AI Systems via System-level DPO

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators