Articles from Source: The-New-Stack

The “silent hallucination” loop: how our autonomous data pipeline poisoned its own vector store

2026-07-09 13:00
🚨 In a recent article, the author discusses challenges faced with an AI system for a fintech client. Initially, the system worked well, extracting data from PDFs. However, it soon generated inaccurate chatbot responses, citing outdated financial data and incorrect revenue attribution. The root cause was a faulty data ingestion process that mismanaged probabilistic outputs, leading to "hallucinations" in the database. Despite implementing a validation layer, issues persisted due to reliance on...
Source: The New Stack
Emmanuel Akita

“The Switzerland of AI”: OpenClaw becomes a non-profit foundation

2026-07-09 12:44
🚀 OpenClaw has evolved into a non-profit foundation, marking a significant milestone in the open-source AI landscape. The initiative aims to empower users by allowing them to run AI agents locally on their machines, enhancing personal automation across various applications like WhatsApp and Slack. Chaired by Dave Morin, the foundation seeks to prioritize user control over AI, distancing it from corporate interests. For more insights, check out the full article! #OpenClaw #AI #OpenSource...
Source: The New Stack
Paul Sawers

“Opus-class, but faster”: What Elon Musk says about beating Anthropic

2026-07-08 18:59
🚀 Exciting news from SpaceXAI CEO Elon Musk! Grok 4.5 will be publicly launched tomorrow. This new model is described as "Opus-class," but faster and more cost-effective. It benefits from a powerful 1.5-trillion-parameter engine and specialized training data from the Cursor platform. Musk aims to compete with Anthropic's Claude Opus amidst recent challenges faced by its Fable 5 model. Stay tuned for more updates! #Grok45 #AI #ElonMusk #SpaceX #TechNews
Source: The New Stack
Amanda Caswell

JetBrains’ next move isn’t a better IDE — it’s a governance layer over Claude Code, Codex, and Gemini CLI

2026-07-08 17:44
🚀 JetBrains is enhancing its offerings with a new governance layer for AI tools used by engineering teams. This layer, called JetBrains AI for Teams and Organizations, will integrate with existing tools like Claude Code, Codex, and Gemini CLI, without requiring teams to switch vendors. The rollout will begin gradually in July and August, aiming to provide shared context, process reuse, and cost controls across tools. For more insights, read the full article. #JetBrains #AI #Engineering...
Source: The New Stack
Paul Sawers

Meta says it caught OpenAI. One thing is missing.

2026-07-08 17:34
🚀 Meta's AI developments are in the spotlight as Alexandr Wang claims their model, "Watermelon," has matched OpenAI's GPT-5.5 on key benchmarks. However, details are sparse, and no specific benchmarks were provided for verification. This leaves the claim unsubstantiated. With recent layoffs affecting morale, the timing of these announcements raises questions. #Meta #AI #OpenAI #TechNews #Innovation
Source: The New Stack
Janakiram MSV

OpenAI’s own safety card says GPT-5.6 has a lying problem

2026-07-08 17:08
🚀 OpenAI has announced the public launch of GPT-5.6 Sol, Terra, and Luna this Thursday. This rollout follows a period of limited access to select partners. 🔍 The models aim to cater to different needs: Sol is for reasoning, Terra focuses on daily workloads, and Luna is designed for high-volume, low-cost inference. 💰 Terra stands out for its cost-effectiveness, offering performance similar to GPT-5.5 at reduced prices. This has generated positive discussions among developers regarding economic...
Source: The New Stack
Amanda Caswell

Most enterprises will hand root cause analysis to AI agents within two years

2026-07-08 14:00
In the next two years, many enterprises plan to shift root cause analysis to AI agents. This change aims to enhance efficiency in managing complex IT systems. Currently, software engineers spend significant time analyzing logs and data to identify issues. However, this process is becoming increasingly burdensome. With the rise of generative AI (GenAI), 85% of organizations now utilize it for observability, with expectations for that number to reach 98% soon. GenAI allows for autonomous...
Source: The New Stack
TNS Staff

“Nature is the most computationally efficient system we know”: How Refiant used swarm optimization to build a 10-million-token AI model

2026-07-08 13:00
Refiant has launched its 10-million-token AI model, Protea, using swarm optimization inspired by nature. 🌿 This approach mimics how ant colonies and other creatures efficiently solve problems, aiming to enhance AI's computational efficiency. Co-founder Dr. Viroshan Naicker emphasizes the need for more than just larger context windows, as many models struggle with memory limitations. The use of nature-inspired algorithms could lead to breakthroughs in AI inference and reduce common issues like...
Source: The New Stack
Adrian Bridgwater

Entire is building a Git network for agents

2026-07-08 13:00
🚀 Thomas Dohmke, former GitHub CEO, is launching a distributed Git network with his startup, Entire. This aims to manage AI coding agents more efficiently and reduce strain on central servers. Entire's system allows developers to mirror GitHub repositories easily, enhancing performance and minimizing rate limits. The focus is on creating a decentralized network for faster operations. Currently in the U.S., EU, and Australia, Entire plans to expand further. They have developed a new backend...
Source: The New Stack
Frederic Lardinois

Coinbase runs 1,200 agents and just slashed its AI bill in half

2026-07-07 20:55
Coinbase CEO Brian Armstrong and Vercel CEO Guillermo Rauch are shifting away from single AI providers. Both leaders are designing systems to utilize multiple models, capitalizing on improved open-weight alternatives and cost efficiency. Armstrong noted that Coinbase halved its AI expenses while usage grew, thanks to strategies like defaulting to lower-cost models and task-based routing. Rauch emphasized the obsolescence of single-lab partnerships, highlighting a trend towards more flexible,...
Source: The New Stack
Amanda Caswell

Watch AWS engineers troubleshoot agentic AI with OpenTelemetry and OpenSearch

2026-07-07 20:29
🚀 AWS engineers tackle the challenge of agentic AI troubleshooting using OpenTelemetry and OpenSearch. As organizations seek better insights into system performance, traditional telemetry methods struggle with AI's complexity. OpenTelemetry offers a unified context, while OpenSearch serves as a key retrieval tool for AI agents. 💡 Join the live webinar on July 22 to watch a troubleshooting simulation and learn about the Agent Health framework, which helps ensure reliable agent behavior before...
Source: The New Stack
Jennifer Riggins

Vercel acquires Better Auth to give AI agents their own identity

2026-07-07 20:13
🚀 Vercel has acquired Better Auth to enhance AI agent identity. AI agents perform tasks like code reviews and deployments but currently share the same identity as users. This acquisition aims to address that issue. Better Auth is known for its TypeScript authentication framework, which has millions of downloads weekly. The team will focus on developing Agent Auth, giving AI agents distinct identities with specific permissions. This move reflects a broader trend in the tech industry toward...
Source: The New Stack
Paul Sawers

Anthropic gives Claude subscribers five more days with Fable 5

2026-07-07 18:45
📢 Anthropic has extended access to its Fable model for Claude subscribers until July 12. Originally set to transition to a pay-per-token model today, users now have five extra days to utilize Fable. This extension allows subscribers to maximize their usage, especially for ongoing projects. Fable was initially available from June 9 to June 22 but faced interruptions. Now, users can enjoy the model without additional costs until the new deadline. #Anthropic #Claude #Fable5 #AI #TechNews
Source: The New Stack
Frederic Lardinois

Anthropic’s Claude Cowork now keeps working when you close your laptop

2026-07-07 16:00
🚀 Anthropic has updated Claude Cowork to now operate on web and mobile, allowing it to run tasks even when your laptop is closed. This means better flexibility for knowledge workers. 📱 Subscribers to the Max plans have early beta access, with broader availability expected soon. 🔗 The update integrates Claude Chat and Cowork into a single interface, simplifying task delegation for users. 💡 This change aims to encourage more experimentation with Cowork for automating workflows. #Anthropic...
Source: The New Stack
Frederic Lardinois

The organizational iceberg: the invisible data breaking your AI agents

2026-07-07 14:00
When building data platforms for large organizations, key decisions often overlook "invisible data." This includes exceptions, undocumented knowledge, and context crucial for operations but rarely recorded. A recent audit revealed that many datasets were orphaned or redundant, leading to inconsistent decision-making. AI agents, while effective in processing data, struggle with understanding the reasoning behind decisions. Understanding this gap is vital, as traditional systems often lose...
Source: The New Stack
Nitin Singhal

How to kill the code review

2026-07-07 13:00
Traditional code review faces challenges in an AI-driven software development lifecycle. As AI tools accelerate coding, the effectiveness of code reviews is diminishing. Data shows a significant rise in code churn, incidents per pull request (PR), and review times. Notably, 31% more PRs are being merged without any review. While code review historically supported team alignment and standards-checking, these roles require different solutions. AI may assist with standards, but human input...
Source: The New Stack
Ankit Jain

A new study just debunked the biggest fear about AI and open source

2026-07-06 21:26
A recent study from Peking University challenges fears about AI's impact on open source projects. Researchers analyzed 1,888 GitHub repositories using AI coding agents and found that newcomer participation remained steady, even as code complexity increased. While cyclomatic complexity went up by 3-4% and cognitive complexity jumped 11% in Python, these changes did not deter new contributors. Interestingly, the active contributor base continued to grow. However, the study focused primarily on...
Source: The New Stack
Amanda Caswell

Microsoft, Google and Cloudflare just made 2029 the new quantum deadline

2026-07-06 20:20
🚀 Microsoft, Google, and Cloudflare have announced a significant shift in the timeline for quantum-safe technologies, moving the deadline to 2029. This change responds to government directives urging organizations to adopt post-quantum cryptography by 2030. Mark Russinovich from Microsoft emphasizes the urgent need for early action to mitigate risks associated with quantum computing advancements. The transition to quantum-safe cryptography is a multi-year effort, and planning early is crucial...
Source: The New Stack
Adrian Bridgwater

Getting Claude Code to grunt in Caveman-speak might not save as many tokens as you think

2026-07-06 18:28
Developers are increasingly focused on the costs of AI coding tools. GitHub's Copilot now uses usage-based billing, while some startups report significant savings by switching AI models. 💻💰 To reduce token consumption, many are adopting a "caveman mode" for AI responses—short and direct answers with minimal filler. This approach aims to cut costs and improve efficiency. 🔍💡 For instance, Elastic's tool achieved a 63.6% reduction in tokens used, highlighting the financial impact of unnecessary...
Source: The New Stack
Paul Sawers

Why most AI projects fail: It’s infrastructure and people

2026-07-06 18:19
🚀 AI projects often struggle to deliver business results, with studies showing a high failure rate—95% according to MIT NANDA. The main issues? Inadequate data infrastructure and a lack of personnel to manage production applications. Prototyping environments typically don’t meet the needs of large enterprises, especially in regulated industries. Flexibility and security are crucial for moving from prototype to production. Organizations must ensure their data management aligns with production...
Source: The New Stack
Meredith Shubel

Palantir’s Alex Karp and Mistral’s Arthur Mensch agree: AI lock-in is coming for enterprises

2026-07-06 17:17
🚀 Palantir CEO Alex Karp recently discussed a partnership with Nvidia on CNBC, emphasizing concerns about the AI model industry. He criticized companies like OpenAI for overcharging and exploiting proprietary data. 📈 Similarly, Mistral CEO Arthur Mensch spoke on LinkedIn about the risks of closed AI providers gaining power over enterprises. Both leaders advocate for open-weight models and proprietary systems to counteract this trend. Their insights highlight the growing concern of AI lock-in...
Source: The New Stack
Amanda Caswell

Andrej Karpathy, Google and Garry Tan agree Markdown is the answer, but they’re not solving the same problem

2026-07-06 15:13
Andrej Karpathy's "LLM Wiki" aims to build personal knowledge bases using Markdown files for AI agents. 📚 Google has introduced the Open Knowledge Format to standardize organizational knowledge in Markdown, while Garry Tan's gstack provides a unique setup for engineering teams, also utilizing Markdown. 🔧 Each approach addresses different needs but highlights the growing importance of Markdown as a resource for AI. The focus is shifting from AI models to the knowledge captured in Markdown...
Source: The New Stack
Janakiram MSV

The code review bug hunt is dead. Here’s what developers get wrong.

2026-07-06 12:00
The code review process is essential for quality assurance in software development. It aims to catch bugs early and promote mentorship among developers. However, many companies fail to define clear outcomes for these reviews. Mark Dominus emphasizes that code reviews should focus on maintainability rather than solely bug detection. He warns against relying on reviews to find bugs, as this can lead to misunderstandings and frustrations. Dominus suggests that project leaders should guide the...
Source: The New Stack
Adrian Bridgwater

Microsoft, AWS and Anthropic are spending billions — and not on better models

2026-07-05 15:00
📢 Microsoft has announced the formation of the Microsoft Frontier Company, aimed at embedding 6,000 experts within customer organizations to enhance AI deployment. This initiative is backed by a $2.5 billion investment. AWS also recently committed $1 billion to a similar engineering organization, indicating a shift in the AI industry towards engineering resources rather than just model development. Microsoft's Frontier Transformation will focus on industry expertise and enterprise AI...
Source: The New Stack
Janakiram MSV

10 moments that defined AI’s turbulent first half of 2026

2026-07-05 14:00
🚀 AI continues to dominate headlines in 2026, showcasing significant developments. The U.S. Commerce Department recently ordered Anthropic to take down its models but lifted the ban shortly after. The Pentagon also clashed with the company over military access to its AI models. Notably, AI firms like Anthropic and OpenAI are eyeing IPOs worth over $800 billion, pushing for infrastructure enhancements amid rapid model releases. Key moments include President Trump's executive order for AI...
Source: The New Stack
Darryl K. Taft

The AI revolution will not be televised — it’ll be quantized

2026-07-05 13:00
The AI revolution is being driven by quantization, a process that compresses AI model weights for efficiency and affordability. In China, the commitment to open weights allows developers to customize and run models like Qwen and DeepSeek locally, enhancing their control over AI capabilities. These models serve as valuable tools in software development, aiding in tasks such as test generation and debugging, though they still require human oversight. #AIRevolution #Quantization...
Source: The New Stack
Adrian Bridgwater

Why cheaper models alone won’t save your AI budget

2026-07-04 12:02
AI budget management is evolving as token consumption rises. 🤖 Engineers face challenges with high token usage per agent operation, leading to increasing costs. For instance, a simple task can consume up to 200,000 tokens when using multiple agents. To address this, teams are focusing on reducing unnecessary token transfers and optimizing workflows. Strategies include compressing context and routing tasks to more cost-effective models. Learn more about these solutions! 💡 #AIBudget...
Source: The New Stack
Amanda Caswell

Apple just turned Safari into something AI agents can control

2026-07-03 17:19
🚀 Apple recently launched Safari Technology Preview 247, featuring a built-in Model Context Protocol (MCP) server. This allows AI agents to interact directly with the Safari browser. 🛠️ Developers can now utilize 16 tools to capture screenshots, inspect elements, and run accessibility checks without leaving the terminal. 🔒 Importantly, this server operates locally, ensuring user privacy by not accessing personal data or browsing history. This update signifies a shift in how browsers can...
Source: The New Stack
Amanda Caswell

The $1.3 million theft that exposed AI’s blind spot

2026-07-02 21:05
A recent cargo theft in Chicago has highlighted a new vulnerability in AI infrastructure: the physical supply chain. 🏗️🚚 Two trailers containing $1.3 million worth of data center equipment and copper wire were stolen from different locations. This incident underscores that as AI demand grows, the risk of theft of essential hardware is increasing. Supply chain delays can impact the entire deployment of AI systems, making this an emerging concern for the industry. 📈🔒 #AI #CyberSecurity...
Source: The New Stack
Amanda Caswell

Microsoft just admitted its biggest AI mistake — and spent $2.5 billion fixing it

2026-07-02 21:03
Microsoft has announced a significant shift in its AI strategy, launching a new $2.5 billion initiative aimed at helping businesses customize their AI solutions. This move comes as the company acknowledges past mistakes, particularly in relying solely on OpenAI models. The new Microsoft Frontier Company will assist clients in integrating various AI tools from different providers, prioritizing flexibility and effectiveness. With this approach, Microsoft aims to enhance AI deployments tailored...
Source: The New Stack
Amanda Caswell

“AI contributions are demoralizing”: Godot bans coding agents to save its mentoring model

2026-07-02 20:31
Godot Engine is updating its contribution policy to restrict AI-generated code in its repositories. This decision follows extensive discussions within the Godot Foundation, as maintainers face difficulties managing the influx of AI-authored pull requests. The Foundation emphasizes that code reviews are essential for mentoring future contributors, which is compromised when dealing with AI. While AI use is still allowed for simple tasks like code completion, new contributors will need explicit...
Source: The New Stack
Paul Sawers

What comes after attention? This startup says it already knows.

2026-07-02 17:00
🚀 Subquadratic has launched its SubQ 1.1 Small model, which utilizes a unique sparse-attention mechanism to handle a 12-million token context window efficiently. Despite early skepticism, the company has released model benchmarks and is collaborating with design partners. Co-founder Alex Whedon emphasizes that their focus extends beyond sparse attention models to innovative architectures. The SubQ 1.1 shows strong results in long-context retrieval tasks, achieving near-perfect scores on...
Source: The New Stack
Frederic Lardinois

Your social login buttons run on third-party cookies. FedCM doesn’t.

2026-07-02 14:00
🚀 Social login options like "Sign in with Google" and "Continue with Apple" have simplified user onboarding for over a decade. However, they depend on third-party cookies, which privacy regulations are challenging. 🍪 Browsers like Safari and Firefox have already blocked these cookies by default, leaving many users in a cookieless environment. 🔄 FedCM (Federated Credential Management) is emerging as a solution, enabling federated logins without cross-site tracking. This new API allows browsers...
Source: The New Stack
Jeff Hickman

Why traditional CI/CD fails for LLMs (and the release gates we built to fix it)

2026-07-02 13:00
Traditional CI/CD gates often fall short for production AI systems, especially with LLMs. This article outlines a new release-gating approach that includes baseline evaluations, drift detection, and shadow validation. The goal is to catch silent AI regressions before they impact users. Unlike conventional software, LLMs are probabilistic, requiring more nuanced release checks to ensure acceptable behavior. Learn more about the challenges and solutions for deploying LLMs effectively. 🤖📊🔍...
Source: The New Stack
Freddy Daniel Alvarez Pinto

OpenClaw’s new app doesn’t run AI on your phone. That’s the whole point.

2026-07-01 21:00
🚀 OpenClaw has launched its new iOS and Android apps, allowing users to connect directly with their personal AI agents without relying on Telegram or WhatsApp. 📱 The app functions as a remote control, enabling communication with an AI that runs elsewhere, rather than on the phone itself. This design enhances performance and usability. 🔍 Similar approaches are seen in platforms like Anthropic’s Claude and OpenAI’s Codex, indicating a shift in how mobile apps are developed. #OpenClaw #AI...
Source: The New Stack
Amanda Caswell

Cordyceps flaw pattern is more proof CI/CD is part of the attack surface

2026-07-01 20:29
🔍 On June 24, Novee Security revealed a CI/CD vulnerability named "Cordyceps," affecting organizations like Microsoft and Google. This flaw allows unauthorized GitHub users to hijack workflows, compromising open-source supply chains. Out of 30,000 scanned repositories, 654 were flagged, with 300 confirmed exploitable. Developers often neglect CI/CD pipelines as security risks, leading to potential threats. Security scanners also struggle to identify such nuanced vulnerabilities....
Source: The New Stack
Meredith Shubel

Cloudflare wants to build the economic layer of the AI web

2026-07-01 20:10
Cloudflare is evolving to support publishers in the AI-driven web landscape. They recently announced updates including crawler classifications and analytics tools to help publishers navigate changes in traffic caused by AI summaries. With their new Pay Per Use model, publishers can earn revenue when their content is featured in AI-generated answers. This approach aims to create a more sustainable economic system for content creators. #Cloudflare #AIWeb #DigitalPublishing #ContentEconomy...
Source: The New Stack
Amanda Caswell

“You Only Compute Once”: How Clockwork wants to put an end to AI training restarts

2026-07-01 17:30
Clockwork aims to revolutionize AI training with its new solution, TorchPass, designed to minimize disruptions caused by GPU failures. Traditionally, failures require rolling back to the last checkpoint, which is time-consuming and costly. With TorchPass, training jobs can seamlessly transition to a backup GPU, recovering in minutes without losing progress. The launch of the YOCO Guarantee ensures that 90% of failures will be resolved without checkpoint rollbacks. If not, customers receive a...
Source: The New Stack
Frederic Lardinois

The call is coming from inside your pipeline: the anatomy of a Codecov attack

2026-07-01 14:00
In January 2021, a single line of code was added to a widely used bash script, leading to a major security breach. This script sent sensitive environment variables to an unknown IP address for 61 days before being detected. The incident highlights a systemic issue in modern software security. As new tools and integrations emerge, the pipeline becomes the new perimeter that must be secured. Organizations must rethink security measures to protect against such vulnerabilities. 🔒💻 #CyberSecurity...
Source: The New Stack
Zeen Rachidi

How Anthropic is bringing Fable 5 back — and when it’ll cost you

2026-07-01 04:12
🚨 Exciting update from Anthropic! Following the U.S. government's lift of export controls, Fable 5 will return on July 1, available globally across various plans. However, users will initially have a limit of 50% on usage until July 7. After that, access will require usage credits. Enterprise users will not have Fable 5 included in their regular allowance and will be billed through usage credits immediately. Anthropic also shared insights on the recent suspension of Fable 5, revealing that a...
Source: The New Stack
Frederic Lardinois