Skip to content

Question about synthetic preference generation for continual RLHF #211

Description

@aus10mccaffs2026

Hi, I came across AIF-Gen and the CPPO model series on your Hugging Face org while researching how people run preference training on open-weight models. I'm researching tooling for this kind of work, so to be upfront, nothing competitive, just trying to learn from teams doing it hands-on. I'm curious what pushed you toward synthetic preference generation, and how you validate that models trained on generated preferences actually pick up what was intended.

Open to a quick chat? Happy to do email, I'm at austin@aureliusaligned.ai

Thanks either way,
Austin

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Fields

    No fields configured for issues without a type.

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions