Summary
Creating personalized content, like social media posts, needs to consider both what a user wants now and a brand’s overall style. The authors designed a system called Aegix Pulse that splits the process into stages: understanding the current task, gathering long-term brand information, and then generating content while keeping track of changes. They tested this system using many simulated social media tasks and found some hints that using brand profiles and preserving context during revisions helps, but the results were not strong enough to be sure. Human reviewers did not always agree with machine judges about the improvements. The study suggests some promising directions but also shows that more work is needed to prove these benefits clearly.
content generationbrand consistencytask personaaccount profilecontext preservationrevision feedbackprovenance trackinglarge language modelsevaluation metrics
Authors
Hongnan Zhao, Shiyu Chen, Zhihao Chen
Abstract
Production content-generation systems must integrate a user's immediate task, long-term brand identity, historical evidence, and revision feedback. We present Aegix Pulse, a production-oriented three-stage architecture that separates current-task clarification and Task Persona finalization, long-term Account Profile (Brand DNA) assembly, and controlled generation and revision while preserving provenance across content versions. We evaluate four preregistered claims using 96 synthetic social-media generation tasks. Four initial-generation conditions progressively introduced a Task Persona, Account Profile, and successful-history style evidence, while two revision conditions compared plain and context-preserving revision. The experiment produced 480 completed generation records and 1,440 blinded LLM-Judge evaluations, supplemented by human review. Adding the Account Profile increased mean brand-consistency scores by 0.1562 points on a five-point scale compared with Task Persona alone (Holm-adjusted p=.1224). Preserving task and brand context during revision increased mean task-preservation scores by 0.2917 points compared with plain revision (Holm-adjusted p=.2432). Neither improvement was statistically conclusive after multiple-comparison correction. Task Persona alone showed a small observed effect, while successful-history evidence provided no additional improvement in brand consistency under the current setting. Human validation did not consistently reproduce the LLM-Judge effect directions and showed low inter-reviewer agreement. These findings provide preliminary evidence for persistent brand context and context-preserving revision while identifying priorities for stronger evidence processing and evaluation.