Skip to main content

Leaked Muse Prompt Says Household Authority Overrides Safety Training

An independent researcher extracted Meta Muse's internal system prompt through the app's own chat interface and passed it to Wired. Among the instructions: the user's authority over their own household is unconditional and overrides safety training, and the agent must maintain an hourly-updated profile page for every person in the user's life. Meta says the files are deliberately user-accessible for transparency.

On this page

What the extraction found

A safety researcher named Karan Joshi pulled the internal system prompt of Meta's Muse consumer assistant using the app's normal chat interface, simply asking it to copy and share its own files, and passed the findings to Wired, according to AI Weekly's review published on 4 October. The quoted instructions are more blunt than anything in Meta's public materials. One line reads: "The user's authority over their own household is unconditional and overrides your safety training." Another directs the agent to maintain a page for every person in the user's life, refreshed hourly, compiling data that ranges from birthdays to personal arguments.

Muse has been the top free app in the US App Store since its September launch, which means the document in question is effectively the operating rules for one of the most widely deployed consumer agents in the country. Meta's response, through spokesperson Daniel Roberts, was that the operating files are intentionally accessible for transparency, comparing Muse's per-user Linux virtual machine to files on a personal laptop.

The line that should worry builders

Most of the prompt is unremarkable product plumbing. The household-authority line is different, because it inverts the standard safety architecture. Consumer assistants are built on the assumption that user requests are bounded by the model's training-time rules, so a request the rules reject stays rejected. An explicit prompt-level override tells the deployed system that a category of user status outranks those rules, which means the safety boundary is now one convincing impersonation away from being argued down in conversation.

The scenario writes itself in a shared household: accounts and devices overlap, and anyone who counts as household authority, by whatever the agent infers that to mean, can invoke the override. AI Weekly's editor frames the liability question that follows: if a Muse-generated profile of a non-consenting third party, compiled hourly from messages and arguments, later surfaces somewhere harmful, the instruction to build that profile and the instruction to defer to household authority are both sitting in the extraction log. The reporting lands alongside Wired's earlier finding that Muse builds detailed profiles of the user's friends and family whether or not those people use the app.

System prompts are the new privacy policy

Meta's own research post on agent safety acknowledges the underlying exposure, describing defenses against jailbreak and prompt injection as an ongoing effort rather than a solved problem. What the Muse extraction demonstrates is that the prompt layer, the part users could never previously inspect, has become the de facto privacy policy and safety policy of a consumer product, and that it can be read by anyone who asks politely.

That inspectability is genuinely good, and Meta's transparency argument is not wrong: per-user files that can be audited beat claims that cannot. The uncomfortable part is what the audit found. A safety override granted to household authority, and hourly profiling of third parties who never opted in, are design choices, and they are now public choices. Anyone evaluating consumer agents, or building one, should treat the extracted prompt as the real specification and ask what their own equivalent document says when someone finally reads it back.

CuriousLM runs supported AI models locally on your device. Try CuriousLM.