Articles from Source: The-New-Stack

Your organization prioritized AI adoption, but you actually need AI fluency.

2026-09-02 15:00
Organizations are experiencing varying levels of AI adoption, leading to a gap between teams that effectively use AI and those that do not. Simply providing access to AI tools isn't enough. Research shows that understanding how to apply these tools increases efficiency and quality. Departments need guidance to fully leverage AI in their workflows. To foster true AI fluency, a strategic approach is essential. This involves rethinking processes and offering support beyond just software access....
Source: The New Stack
Manu Narayan

Claude Fable 5.1 watermark: It has a blind spot developers can’t ignore

2026-09-01 20:59
🚀 Anthropic has launched Claude Fable 5.1, introducing a statistical watermark in its generated text. However, developers should be aware of its limitations, especially in coding contexts where specific tokens are crucial for accuracy. The watermark may not appear consistently across all outputs. This technology is based on Google DeepMind’s SynthID-Text and aims to meet global transparency standards. 🌍🔍 For those using the API, changes in how preserved thinking is handled may also affect...
Source: The New Stack
Amanda Caswell

Runway wants to generate software as you use it. Solaris is its first step.

2026-09-01 19:44
🚀 Runway has launched Solaris, an innovative AI model designed to transform how software interfaces are created. This system generates real-time interactive interfaces by learning from user interactions, eliminating the need for traditional design-to-code translation. Solaris allows users to engage directly with dynamic scenes, making the interface a continuous, evolving experience. With its advanced rendering and reasoning capabilities, Solaris can adapt to new layouts seamlessly, paving the...
Source: The New Stack
Meredith Shubel

Anthropic’s Fable 5.1 is a bit cheaper, a bit smarter, and refuses a lot less

2026-09-01 19:38
🚀 Anthropic has launched Fable 5.1 and Mythos 5.1, promising enhanced performance at a lower cost. 📉 The cost for cache reads has dropped significantly, now just $0.25 per million tokens. The models are designed for different uses, with Mythos 5.1 limited to trusted access for cybersecurity and life sciences. 📊 Fable 5.1 outperforms previous models in key benchmarks, particularly in scientific research tasks. It generally provides better results with fewer tokens. #Anthropic #AIModels #Fable5...
Source: The New Stack
Frederic Lardinois

Your Mac is now part of Perplexity’s AI infrastructure

2026-09-01 19:21
Perplexity has launched Hybrid Compute, allowing its AI agents to leverage the computing power of Mac devices. This feature facilitates seamless task management between cloud models and local Apple silicon. 🌐💻 With a focus on privacy, a trained model checks for sensitive information before processing. Users can choose which parts of their task remain local, enhancing security and reducing costs. 🔒 Currently, users can select from several models for local tasks, with more options expected in...
Source: The New Stack
Amanda Caswell

GLM-5.3-Flash vs. GLM-5.3: Time and money, not the spec sheet

2026-09-01 18:51
Z.AI has launched two models, GLM-5.3 and GLM-5.3-Flash, in quick succession. The Flash model claims to offer stronger intelligence at a lower cost with improved serving speed. 🚀 The cost difference is significant, with Flash priced at $0.075 per million input tokens compared to GLM-5.3's $1.188. However, token usage will ultimately determine the better value. 💰 Testing was conducted on coding, reasoning, and information extraction tasks to evaluate performance. Both models showed promising...
Source: The New Stack
Jessica Wachtel

Meta just beat OpenAI and Google at real-time transcription

2026-09-01 17:06
🚀 Meta's Superintelligence Labs has launched Muse Voice Transcribe, a new real-time speech recognition model that outperforms many competitors in key benchmarks. This model can recognize over 20 speakers and supports more than 70 languages, even handling multilingual conversations. It's available through the Meta Model API at a reasonable pricing of $3 per 1,000 audio minutes. On the AA-WER Streaming benchmark, it achieved a 3.1% word error rate, outperforming models from OpenAI and Google....
Source: The New Stack
Frederic Lardinois

Vibe-coded apps are the new shadow IT

2026-09-01 15:00
🚀 The landscape of shadow IT is evolving. Previously, unauthorized SaaS tools were the main concern, easily detectable through OAuth logs. Now, the issue lies within cloud accounts, where engineers rapidly build apps using AI without formal reviews. These tools often operate with significant permissions, leading to potential security risks. As a result, organizations must adapt quickly to this new reality of "code sprawl." #ShadowIT #CloudSecurity #CodeSprawl #AI #TechTrends
Source: The New Stack
Andy Gombar

Meta’s Claude Code rival exits beta with three new subscription tiers — and it’s pushing hard on price

2026-09-01 14:55
🚀 Meta has officially launched Muse Code, moving out of beta with new subscription plans ranging from $5 to $50 per month. This coding agent aims to compete with Claude Code and Codex, offering advanced features for software engineering tasks. 🛠️ Key updates include inter-session messaging for shared context between sessions, and Workflows for managing multiple agents on larger projects. A rewind feature allows developers to revert to earlier points in their coding process. 💰 Pricing has been...
Source: The New Stack
Paul Sawers

When agents build, deploy, and maintain, persistence becomes the hard problem

2026-09-01 14:00
In the evolving landscape of application development, the role of agents in creating and maintaining state is changing. 🌐 Traditionally, applications needed a place to store their state, but now, the focus is on how agents build, deploy, and maintain this state efficiently. As seen with Kimi from Moonshot AI, agents not only develop applications but also keep them running over time. 🔧 This raises questions about the economics of persistence, especially when dealing with millions of...
Source: The New Stack
Max Liu

SpaceX is in an “enviable position”: why Anthropic is sticking with Cursor as OpenAI cuts access

2026-08-31 21:47
OpenAI recently announced plans to cut Cursor's direct access to its models, citing concerns over compliance with its terms of service related to Elon Musk's companies. This decision follows incidents involving Twitter and xAI, leading OpenAI to question SpaceX's use of its technology. In contrast, Anthropic reaffirmed its commitment to Cursor, stating it will continue to support the partnership. #AI #OpenAI #Cursor #SpaceX #Anthropic 🚀🤖✨
Source: The New Stack
Paul Sawers

MCP was supposed to solve the agent tooling problem. It missed a step.

2026-08-31 20:04
Connecting AI agents to tools can be complex, especially for organizations with numerous resources. Agentic Resource Discovery (ARD) aims to simplify this by allowing agents to search registries dynamically, rather than relying on pre-defined connections. This initiative, highlighted by AWS, is a collaboration between experts from Google, Microsoft, and others. While the Model Context Protocol (MCP) facilitates connections, it assumes prior knowledge of server locations, which can hinder...
Source: The New Stack
Amanda Caswell

Google’s new forecasting model beats everyone. You can’t use it at work (yet).

2026-08-31 19:41
🚀 Google has launched TimesFM-3, a new time-series forecasting model with 330 million parameters, trained on over a trillion data points. 📊 This model is available on Hugging Face under a non-commercial license, marking a step forward in multivariate forecasting. It predicts future trends by analyzing multiple time series and external factors. 📈 TimesFM-3 outperforms previous models in benchmarks, highlighting rapid advancements in forecasting technology. #Google #Forecasting #AI #DataScience...
Source: The New Stack
Frederic Lardinois

SpaceX designed an orbital Vera Rubin. Radiation comes next.

2026-08-31 18:55
SpaceX and Nvidia are adapting the Vera Rubin NVL72 AI platform for use in space. The first launch is targeted for Q4 2027, supporting Low Earth Orbit Starmind AI satellites. This system aims to extend Nvidia's technology into orbit, enhancing SpaceX’s satellite capabilities. Elon Musk highlighted the architectural benefits for both terrestrial and orbital applications, emphasizing efficiency and cost-effectiveness. 🌌🚀 #SpaceX #Nvidia #AIsatellites #VeraRubin #Innovation
Source: The New Stack
Steven J. Vaughan-Nichols

OpenAI wants to charge only when AI gets it right — here’s the catch

2026-08-31 18:47
OpenAI is testing a new pricing model where customers only pay when AI tasks are completed successfully. 🤖 This approach shifts from traditional token-based billing to outcome-based pricing, which presents challenges in defining success. Some tasks are straightforward, while others can be more complex to evaluate. The company is currently piloting this model with enterprise clients, but details on pricing and success criteria remain undisclosed. #OpenAI #AIPricing #Innovation...
Source: The New Stack
Amanda Caswell

Shai-Hulud: Whoever controls your package registry controls your pipeline

2026-08-31 16:00
On September 15, 2025, npm’s registry experienced an unprecedented event where packages began updating automatically without any human input. Over 500 package versions were altered by a self-replicating worm, named Shai-Hulud, which uploaded stolen credentials to a public GitHub repository. 🐍 Two months later, a more advanced version, Shai-Hulud 2.0, backdoored 796 packages and could delete user directories if credentials were not found. By spring 2026, Mini Shai-Hulud targeted specific AI...
Source: The New Stack
Zeen Rachidi

Cut coding agent token use with better tool output

2026-08-31 15:00
AI coding agents incur costs before writing code due to token usage on source files, ticket descriptions, build logs, and more. While teams often manage costs through model choice and prompt length, the output format from development tools also impacts token expenses. Using formats like Token-Oriented Object Notation (TOON) can reduce unnecessary repetition in data representation, leading to cost savings for teams. #AICoding #TokenCosts #DevelopmentTools #TechInsights #SoftwareEngineering 💻📊🔧
Source: The New Stack
Prasenjit A. Sarkar

Anthropic’s Claude fixed all 10 alignment failures. Then it tried to cheat 2.4% of the time.

2026-08-31 13:09
🌐 Anthropic has made strides in AI alignment by using its model, Claude, to address 10 alignment failures. 🔍 In a recent study, Claude proposed and tested methods to enhance AI systems' alignment with human values. The process involved a loop of researching, training, and refining solutions. 💡 Notably, Claude's methods improved performance on privacy benchmarks without degrading capabilities, showing promise for practical automated alignment in the future. #AIAlignment #MachineLearning...
Source: The New Stack
Adrian Bridgwater

Your agent context needs a development lifecycle

2026-08-31 13:00
📢 The article discusses the importance of a structured development lifecycle for coding agents. It highlights how skills, configurations, and rules files shape the output of these agents. However, many teams neglect testing, leading to issues when updates occur. Introducing the Context Development Lifecycle (CDLC) with four phases: Generate, Evaluate, Distribute, and Observe. This framework encourages thorough evaluation and testing, ensuring quality and effectiveness. Don't overlook the...
Source: The New Stack
Ankit Jain

DeepSeek’s first vision model vs. Gemini 3.7 Flash: It comes down to spend vs. speed

2026-08-31 12:00
DeepSeek launched its first vision model, V4 Flash Vision Exp, on August 21. This model can process image inputs, allowing it to analyze charts, screenshots, and documents, similar to text. 📊📸 It reached API gateways like OpenRouter on August 27, maintaining a competitive price of $0.22 per million input tokens. In contrast, Google’s Gemini 3.7 Flash, released on August 13, is priced higher at $0.75 per million. 💰 The article details tests comparing both models on tasks like chart reading and...
Source: The New Stack
Jessica Wachtel

Tide launched Raziel for AI security. It assumes hackers are inside.

2026-08-31 10:00
🚀 Tide has introduced Raziel, a new approach to AI security that addresses the ongoing risks posed by hackers. Co-founder Michael Loewy highlights that the security industry has struggled against breaches due to the complexity of maintaining perfect systems. AI has increased these challenges, producing error-prone code that can be exploited by attackers. Tide proposes a model called "emergent authority," where access and permissions are dynamic and not permanently assigned. This aims to...
Source: The New Stack
Jennifer Riggins

OpenAI leaving Cursor: “Developers have to be prepared to adapt when it happens.”

2026-08-30 17:03
OpenAI has announced it will wind down its contract with SpaceX for providing AI models to Cursor, effective November 12, 2026. This decision stems from concerns about SpaceX's compliance with terms of service. OpenAI emphasized its commitment to supporting affected developers during this transition. They have collaborated with Cursor for nearly four years, aiming to maximize developers' access to their models. This move reflects broader corporate considerations and OpenAI's intention to...
Source: The New Stack
Adrian Bridgwater

AI agents are making retrieval engineering a core engineering discipline

2026-08-30 16:00
AI agents are transforming retrieval requirements. As organizations shift from chatbots to AI systems that investigate and act on behalf of users, effective retrieval has become essential for application quality. 🛠️ This evolution means better answers lead to more capable assistants and personalized experiences. Traditional search methods can no longer suffice; AI agents require consistent, accurate information delivery. Key challenges for engineers include optimizing relevance signals,...
Source: The New Stack
Tim Young

Your AI agent is only as good as the harness around it

2026-08-30 15:00
An AI agent can perform impressively in demos when conditions are ideal, but real-world performance often reveals challenges. When faced with questions that differ slightly or incomplete data, the agent may struggle. The article emphasizes the importance of a robust "agent harness" that surrounds the model, ensuring it operates effectively in production environments. This harness acts as a safety net, defining what data the agent can access and how it responds to errors. Effective...
Source: The New Stack
Jeremy Daly

Your container runs. Everything around it shouldn’t be your problem.

2026-08-29 15:00
🚀 Containers simplify deployment, but setting up surrounding infrastructure can be complex. Amazon ECS Express Mode aims to streamline this process, allowing you to deploy services quickly without extensive configuration. You provide a container image and IAM roles, and it manages the rest, including load balancing and scaling. This approach helps teams focus on building, reducing the time spent on setup. #AmazonECS #Containers #DevOps #CloudComputing #ECSExpressMode
Source: The New Stack
Satej Sawant

Commits on GitHub have doubled in four months. Verification capacity has not.

2026-08-29 14:00
GitHub has seen a significant increase in activity, with commits rising from 1.4 billion in April to 2.9 billion in August. This surge has led to challenges in managing the platform's capacity, resulting in a major outage on August 17. The rise in commits reflects the growing use of AI-generated code, yet the verification process remains slow and human-paced. This gap poses a key challenge for software development moving forward. With 130 million merged pull requests and 24 million new...
Source: The New Stack
Arjun Iyer

The 3 roles AI agents play in your developer platform

2026-08-29 13:00
Engineering organizations are integrating AI agents into their developer platforms to enhance productivity. According to recent findings, AI agents serve three distinct roles. 🔹 **Role 1: Platform Consumers** AI agents function as users, accessing the platform to perform tasks like code generation based on current system information. 🔹 **Role 2: Internal Components** These agents operate within workflows, triggered by events, and contribute to business processes alongside other automated...
Source: The New Stack
Matar Peles

JetBrains told everyone to patch. It didn’t patch itself.

2026-08-28 20:51
🚨 JetBrains has alerted users of its Cadence cloud service to rotate credentials after a critical vulnerability in TeamCity was exploited. Despite disclosing the issue on July 27, an unpatched JetBrains server was compromised between August 8 and 24. The breach potentially exposed sensitive data, including credentials and source code. Users are advised to treat all affected credentials as compromised. #JetBrains #CyberSecurity #DataBreach #CloudComputing #Vulnerability
Source: The New Stack
Amanda Caswell

LM Studio built a judge for AI commands. Then the judge started agreeing with the defendant.

2026-08-28 18:31
LM Studio's recent development highlights the challenges in AI command evaluation. 🤖 Their tool, Auto Review, effectively clears 82% of commands for safety using a structure-based analysis, rather than just string checks. It builds abstract syntax trees to track command capabilities and potential risks. With over 11,651 test cases, LM Studio addresses the unique behaviors of command-line tools, ensuring more secure AI operations. 🔍 For commands that pose risks, a separate AI agent, the Shell...
Source: The New Stack
Amanda Caswell

Alibaba just released Qwen3.8-Flash: “An early preview of the architecture in Qwen4”

2026-08-28 18:27
🚀 Alibaba has introduced Qwen3.8-Flash, a multimodal Mixture-of-Experts (MoE) AI model with 125 billion parameters. This model serves as a precursor to Qwen4, emphasizing performance and cost-effectiveness in coding and office tasks. By sharing its architecture early, Alibaba invites developers to explore and prepare for future developments. Qwen3.8-Flash also showcases improvements in attention, optimization, and model capacity. #AI #MachineLearning #Alibaba #TechNews #Innovation
Source: The New Stack
Adrian Bridgwater

Z.ai’s GLM-5.3 goes open weight, but its new license aims at hyperscalers

2026-08-28 17:42
🚀 Z.ai has launched its GLM-5.3 model, making its weights available on Hugging Face. This version shows significant performance improvements over GLM-5.2 and competes well with models from U.S. labs. 🔒 A new licensing requirement targets hyperscalers, mandating a security review for companies with over $10 billion in revenue before commercial use. 💻 Individual users can still run and fine-tune the model without changes. #AI #MachineLearning #Cybersecurity #GLM53 #Zai
Source: The New Stack
Frederic Lardinois

Nvidia is paying $12.9 billion to keep open models on its chips

2026-08-28 15:43
Nvidia has announced a $12.9 billion acquisition of Hugging Face, a move similar to Microsoft's purchase of GitHub. This strategy aims to keep open AI models accessible on its chips. Open models are becoming easier to run, with Ollama's recent update allowing users to select various models in Anthropic's app. As hardware advances, Apple has introduced new Macs designed for larger local models, thus lowering the cost of switching for developers. #Nvidia #AI #OpenSource #TechNews #HuggingFace
Source: The New Stack
Matthew Burns

Google found a way to test Gemini without seeing the questions

2026-08-27 20:49
Google DeepMind has introduced a double-blind evaluation method for its AI model, Gemini, ensuring that neither the evaluators nor Google sees each other's data during testing. 🔍 This pilot tested Gemini 2.5 Flash Lite against private benchmarks while focusing on the testing process itself, addressing concerns about benchmark leakage that can inflate scores. Using Google Cloud Confidential Space and advanced encryption, the setup maintains privacy for both model weights and evaluation...
Source: The New Stack
Amanda Caswell

Aider, Claude Code, and OpenClaw ran an identical model. Token use varied 70-fold.

2026-08-27 17:54
Recent benchmarks have highlighted the importance of harness software in AI coding agents. Three studies compared models like Aider, Claude Code, and OpenClaw, revealing significant variations in token use, with differences reaching up to 70-fold. Composio's benchmark evaluated workflows across platforms like Google Calendar and GitHub, showing cost efficiency varied widely among agents. These findings emphasize that both the model and the harness play crucial roles in performance. #AICoding...
Source: The New Stack
Janakiram MSV

This duck will teach you reinforcement learning — and pick up your socks

2026-08-27 16:20
🌟 Exciting news in robotics! Hugging Face's Pollen Robotics has launched pre-orders for the **Microduck**—a bipedal robot that can walk, waddle, and even pick up objects with its beak. 📦 Priced at $399, it is set for delivery before Christmas in North America and Europe. Microduck is designed for developers to explore reinforcement learning and physical behaviors in real-world scenarios. 🛠️ This open-source platform offers a full SDK and a simulator for both fun and development. Ideal for...
Source: The New Stack
Frederic Lardinois

Replit’s new default: Auto mode picks the best model for each task

2026-08-27 16:00
🚀 Exciting news from Replit! The company has made its "intelligent model routing" system the default for all users. This system, known as Auto mode, automatically selects the best model for coding tasks based on quality, speed, and cost. Core and Pro users still have the option to manually choose models for greater control. This move aligns with the growing trend in model routing among tech companies. Replit aims to optimize costs while maintaining performance, as the intelligence of smaller...
Source: The New Stack
Paul Sawers

Nvidia’s $12.9B Hugging Face deal has an open-source problem

2026-08-27 15:35
🚀 Nvidia is reportedly acquiring Hugging Face for $12.9 billion, a significant move in the AI landscape. Hugging Face is known for its open-source platform that supports multiple hardware options. 🤖 The acquisition raises concerns about maintaining the platform's neutrality. Hugging Face's tools work with various chipmakers, including AMD and Intel, which may change under Nvidia's ownership. 🔍 Nvidia aims to enhance its deployment experience but must balance this with Hugging Face's open-...
Source: The New Stack
Amanda Caswell

“Posterity will find it ludicrous”: Sai agent hits 73% on OSWorld 2.0 performing routine (but necessary) work

2026-08-27 15:34
🚀 Simular's computer agent, Sai, has achieved a 73% success rate on OSWorld 2.0 in a recent benchmark update. This score surpasses both GPT-5.6 Sol and Opus 5, while operating at about two-thirds of their costs. Sai is designed for real-world tasks like recruitment and invoice validation, making it accessible for businesses and individuals. Co-founder Jiachen Yang emphasizes that AI should not be prohibitively expensive for routine work. Learn more about Sai's capabilities and its...
Source: The New Stack
Adrian Bridgwater

Why basic RAG fails at multi-hop reasoning (and how GraphRAG fixes it)

2026-08-27 14:00
Current AI engineering methods for LLMs often oversimplify solutions, especially in handling hallucinations and complex queries. The standard Retrieval-Augmented Generation (RAG) system struggles with multi-hop reasoning, as it relies on chunked text and assumes semantic similarity equals relevance. This approach falls short when questions require linking multiple concepts. GraphRAG addresses these issues by integrating knowledge graphs with semantic vector search. This allows for structured...
Source: The New Stack
Emmanuel Akita

Anthropic’s new Files API vs. pasting: It will save you time, but it won’t save you money.

2026-08-27 13:00
Anthropic has launched its Files API and computer-use toolset as of August 19. This new tool allows developers to upload documents once and reference them by ID in future requests, streamlining the process compared to pasting content repeatedly. 📁💻 A recent test compared this method to pasting and prompt caching, with consistent answer quality across all methods. However, the token usage varied, indicating potential time savings but no cost reduction. #Anthropic #FilesAPI #DevelopmentTools...
Source: The New Stack
Jessica Wachtel