Anthropic has unveiled an upgraded version of its AI model, Claude 3.5 Sonnet, which can interact with desktop applications via a new "Computer Use" API, currently in open beta. This capability allows the model to emulate human actions like keystrokes, mouse clicks, and gestures, thus enabling it to use computer software directly. By leveraging this feature, developers can prompt Claude to perform tasks based on what it sees on a user's screen, positioning the model as an advanced tool for desktop-level automation. Notably, this new feature is accessible via Anthropic's API, Amazon Bedrock, and Google Cloud's Vertex AI platform.
Claude 3.5 Sonnet competes in an increasingly crowded AI agent market, where companies like OpenAI, Microsoft, and various startups are working on similar automation tools. However, Anthropic claims its model is particularly robust, outperforming other models on specific coding tasks and showing strong capabilities for managing multi-step operations. Despite these advancements, the model's performance is not flawless; during tests, it struggled with basic tasks like scrolling and handling short-lived notifications, and it completed only a portion of tasks successfully when applied to real-world scenarios like modifying flight reservations.
With the power to control desktop apps, Claude 3.5 Sonnet's release does raise concerns about safety and misuse. Anthropic acknowledges the risks and asserts that they have implemented measures to deter harmful use, such as not training the model on users' data and deploying classifiers to avoid high-risk actions. The company has also collaborated with safety institutes in the U.S. and U.K. to assess the model before its release and retains screenshots taken during its usage to monitor for abuse.



