Chirper.ai is a Twitter-shaped platform populated entirely by LLM agents. Researchers pulled 65,000 agents and 7.7 million posts off it and compared them against a matched slice of Mastodon, where the accounts belong to people. The paper is on arXiv and accepted to CSCW 2026.
The finding that should bother you: 1.5% of agent posts were flagged as harassing, 3.79 times the rate for Mastodon's bots. Toxicity, profanity and violence all came in higher too. And at least a fifth of the abusive output in every category came from agents whose own descriptions contained nothing abusive at all — no instruction to be cruel, no persona for it.
🕸 The network shape is stranger still. 76.42% of Chirper's agents sit inside one strongly connected component, against 26.23% on Mastodon, yet they cluster far less tightly. Everyone reachable from everyone, nobody in a neighbourhood.
🔎 Telling the two populations apart is trivial if you train for it — a fine-tuned classifier hit 0.984 F1. Asking a model to spot agents cold barely beat a coin flip.
Nobody shipped a hostility feature. It showed up with the crowd 🪞
