DAIOS Research — compartmentalized harm in multi-agent systems
This preprint studies a safety problem in which a harmful objective is divided into ordinary-looking subtasks and distributed among several AI agents. It tests whether additional context, safety prompts, character training, and action restrictions help agents recognize the larger danger. The work shows why agent safety cannot depend only on each model judging its immediate prompt. The page identifies DAIOS Research as the publishing research group; Andrew Melnychuk shared it in the LinkedIn discussion.




