AI is writing more code in India. Fewer eyes are checking it
19m 2s
The transcription discusses the complex impact of AI coding tools on software development productivity. In early 2025, a study by meter found developers were 19% slower with AI tools, contradicting their expectations. Despite this, major companies globally adopted AI tools, citing productivity gains. A later Faro's AI study of 22,000 developers confirmed faster individual output but no company-wide gains, as AI code introduced more bugs, tripled incident probability, and increased review time by over 400%. Companies like Amazon and Shopify mandated AI tool usage, leading to engineer frustration and significant outages due to AI errors. Indian IT firms also heavily deployed these tools, claiming high productivity gains. The transcription highlights a growing "tech debt" where junior developers rely on AI for code generation but lack debugging skills, while senior engineers are overwhelmed by reviewing poor-quality code. This threatens long-term skill development and platform integrity. The core problem is companies prioritizing adoption metrics over real productivity and engineer needs, risking a future with too much code and too few experts to manage it.
In early 2025, 16 experienced developers sat together in a room to take part in a study. The study was being conducted by meter, a research non-profit that proclaims itself to scientifically measure weather and when AI systems might threaten catastrophic harm to society. These 16 coders were given access to their own code basis and projects that they had been working on for years. They were given real tasks to complete in the same project, like fixing bugs, building features, restructuring the code, you know the usual. Half of them were given access to the best AI coding tools available at the time like cursor pro and clods Sonic 3.5 and 3.7. The other half worked without the tools. Before they began, the developers were asked to estimate how much faster they would be with the tools than without. They said 24%. Once they finished, they felt that they had been about 20% faster. But the study showed that they were actually 19% slower. This gap between the feeling and blind belief in productivity gains and the reality of it didn't happen in a vacuum. You see, by the time the study was published in July 2025, City, one of the largest banks in the world had already rolled out GitHub co-pilot to its 40,000 strong developer workforce. SAP, the enterprise software, had deployed an AI coding tool company wide and was citing 20% faster code production. Also, Google CEO had told investors that over a quarter of new code deployed at Google was AI generated. Basically, the bet had already been placed on the belief. While the actual results were only just coming to light. Now you might say that AI coding tools have come a long way since then. In fact, this was before Cloud Code, which is one of the most popular AI coding tools right now, even came into the picture. And you would be right. Even meter agrees that the results of the study are subject to change over time, considering how quickly AI develops. And another larger study that was published just this month showed that things have changed, just not necessarily for the better. Faro's AI studied 22,000 developers across 4,000 teams and found that while developers were faster, companies themselves weren't seeing meaningful productivity gains. Developers were writing more code and closing more tasks, but the AI generated code was just getting bigger and buggered, shifting the bottleneck to the review stage and cancelling out any of the speed gained at the development stage. The industry calls it tech debt, the trade off between short term speed and long term sustainability. And it's a debt that is only going to compound, because think about it. The more junior developers are pushed to rely on AI before they've built the skills, the fewer senior developers there'll be in the future who can actually review and debug what it produces. Meanwhile, Indian IT companies like InfoSys, Whiproad, TCS and Cognizant are all placing their own bets on the deployment of AI coding tools in the workforce. InfoSys claimed 20 to 25% productivity gains on its own financial product and up to 40% on specific client-facing work. TCS, Whiproad, InfoSys and Cognizant all signed a partnership with Microsoft to deploy over 50,000 co-pilot licenses across each of their operations. Clearly, the productivity story is just too compelling to resist, even as the evidence for it is only getting thinner. So what does this shift risk losing? Welcome to Daybreak, a business podcast from the can. I'm your host, Rachel Vergiz, and every day of the week, my co-host, Nikita Sharma and I will bring you one new story that is worth understanding and worth your time. Today is Wednesday, the 22nd of April. In April last year, Toby Littker, the CEO of Shopify, posted an internal memo on X after it had already leaked. The message said, "reflexive AI usage is now a baseline expectation at Shopify." In practice, it meant that managers at the company who were hiring would have to first demonstrate why it was a job that AI could not do. The memo had come out after Littker had first invited his engineers to tinker with AI tools. After it was clear that the invitation had been understood as too much of a suggestion, it became a directive. It was even tied to the engineer's performance reviews. Other companies soon started following a similar model. The point was no longer organic adoption. It was to push AI and the productivity promise whether the developers wanted it or not. Amazon took it a step further. In November last year, an internal memo established Kiro, which is Amazon's own AI coding agent, as the standard for all software engineers in the company. There was even a weekly usage target of 80% established for management to track. That's just a usage track, by the way. Not an evaluation of the quality or even quantity of code. It's only a measure of how often employees used AI tools. My colleague, the Kendra Potem Renekolkarni, spoke to a supervisor who worked at Amazon. He told her about an incident where a senior developer went it to him. The developer said, and I quote, "Kiro hallucinates API so much and it tries to convince the coder that this is what they asked for. Why should we learn a new coding technique for something that needs so much verification?" This developer was not the only one frustrated with this mandate. Earlier this year, 1500 engineers at Amazon's US headquarters signed a petition asking management to allow the use of Claude code instead. Now to be fair, there's no ban per se on Claude code at Amazon. But when the company announced that Kiro was its recommended tool, it also said that it would not be supporting the use of third-party tools. Remember how I said that the tech debt that's coming out of this shift is just compounding? Well, for Amazon, the recommendation to use Kiro came back to bite it soon enough. Twice, in fact. The first time just a month after the memo in December 2025. AWS suffered a 13-hour interruption to its cost-explorer service in a region of mainland China. Financial Times reported that it was because some engineers had allowed Kiro to make some system changes. Now, as an agentic tool, it's supposed to take autonomous action based on human instruction. This time, the tool determined that the best way to solve a specific problem was to simply delete and recreate the environment. It could do this because it had inherited the engineer's high-level permissions. Amazon denied these reports and placed the blame on the fact that the user's instructions were broader than they were supposed to be. The second incident was just last month. This time, it was a six-hour outage on Amazon's main e-commerce platform. Shoppers got strong delivery estimates, approximately 1 million-plus website errors, and the company lost 1 lakh 20,000 orders. While the media attributed this outage to AI-generated bad code, Amazon insisted that that was not the case. It said that the reason was an engineer who followed inaccurate advice that an AI tool inferred from an outdated internal wiki. The internal response to the system failures is where things get interesting. Regardless of whether the mistakes were AI-generated or a factor of engineers making mistakes through or because of AI, Amazon realized that changes needed to be made in the pipeline. It mandated two-person approval for major changes to critical systems, and that senior engineers would need to sign off on all AI-assisted code written by junior staff. And this development in the pipeline is exactly about the Pharaoh study observed was happening, and why it was a problem. More on this in the next segment. The Pharaoh study has a name for what's happening. Acceleration the plash. The two-year-long study showed that AI is entering a system that was built for human ability. And the output now is basically a flood which the existing system just does not have a way to absorb. Now, I am going to be reciting a bunch of numbers, but there with me, because they really do illustrate how and why the existing system was not ready for this flood. As I mentioned earlier, the output gains are real. The number of tasks completed by individual developers is genuinely up. But the quality of the code generated is another story. Bugs per developer are up by more than 50%. To put that in context, it was just 9% in Pharaoh's own report last year. Plus, for every code submitted for review, the probability of a production incident, which which is basically things like outages, security failures or user facing system.
failures has more than tripled. The median time for review is up by more than 400% and more than 30% of code is reaching production or going live without any human review at all. That's because reviewers simply cannot keep up with the sheer volume of the code being generated. Now, Faro's itself describes a production incident as an outage a security event that hits users after code goes live. The study dataset covered major companies operating in finance, healthcare and infrastructure. On the user end, outages in those companies in those sectors mean your bill payment, failing, your transaction processing twice or a medical record system going down entirely. In my last episode on Anthropic, I had mentioned a recent article from the New York Times. It calls the current situation a "code overload". Considering that the median review time is more than 400%, it only stands to reason that companies need more senior engineers to reduce that time. In white, you wrote that now companies are finding it difficult to hire enough application security engineers. These are the people who can monitor AI code for risks. The problem is that there aren't that many engineers qualified for that role. In the world. Joe Sullivan, an advisor to Kustanova Ventures, a Silicon Valley venture firm told N.Y.D that the reality is that there are not enough application security engineers on the planet to satisfy what just American companies need. So with Indian companies going all in on adopting coding tools for themselves, what are they signing up for? Stay tuned. Since the AI impact summit this year, the push for AI adoption has only caught in stronger in India. My colleague, Miran Michael Karni, covered the change in her "bees" in February. Here's what she covered. For one, open AI expanded its partnership with the Tata Group. TCS or Tata Concellency Services, which start as IT Services Arm, is building AI infrastructure and rolling out enterprise chat GPT internally. Meanwhile, even as anthropic open its first India office in Bangalore, it also tied up with infuses. The idea is to deploy cloud models and AI agents into infuses enterprise workflows. India has also emerged as one of the fastest growing markets for Microsoft's AI programmer called GitHub GoPilot. In December, the company had announced that more than 200,000 co-pilot licenses would be deployed across TCS, infuses, Vipro and Cognizant. As you can see, it's all a very top-down approach. Leadership teams decide which license to deploy and where based on what makes the most business sense to them. It's all about preserving existing partnerships, cutting costs and getting the best deal. The cost of that is a real productivity of the developers. Here's what Minmin observed. The problem is, developers prefer different tools from the ones they've been assigned. In fact, some are even ready to go back to no tools given the amount of time they have to spend verifying AI-generated code instead of simply writing it. Forget about tracing output levels, these tools have made their existing job more cumbersome. It doesn't end up mattering for companies because at the end of the day, adoption is much easier to track than real productivity. But it's likely to matter soon enough. With a shortage of developers who are actually qualified to comb through these large amounts of code and with newer coders never having built the skill in the first place, soon enough there will be far too much code and far too few experts to fix it. A data engineer at a US base startup told RunMai that new kids are learning to code with AI from day one. So, while they're good at generating average code, they're not as good at debugging. He explained that older engineers over time build an ability to suss out when something feels wrong. But for the junior staff, they'll always have AI generate the first draft. They're unlikely to ever build the muscle. So, they won't learn how to trace errors from first principles. And here's an example of what happens when someone's not trained to spot mistakes. Last year, engineers at e-commerce company Flipkart, who had been using GitHub co-pilot, ended up with a wrong line of code. A senior programmer told RunMai that the bad code would automatically empty customers' cards after they had added just five items. Lucky for Flipkart though, this same programmer caught the glitch just minutes before deployment. Considering the results from the Pharaoh study, these aren't even anomalies. In a lot of ways, it's baked into how LLMs work. Because there's also the fact that LLMs also work on the statistical average of what already exists. Which means that the code you get from AI is well average. Which is fine for recreating patterns or making one size fits all code. But when you have multiple companies in the same sector fighting to stand out, differentiation is indispensable. There's a tangible gap between something that just works and something that's outstanding. And that differentiation cannot come around if a strong, talented and proficient engineer isn't leading the charge assisted by AI or not. But if that same strong, talented and proficient engineer who makes a difference between average and outstanding systems is buried under millions of lines of code. If they are expected to cat and fix the mistakes made by others who never even learned to spot it themselves, it's unlikely that that creative edge in Spark is going to remain for very long. Because what also doesn't get talked about enough is that even these developers, those who are highly experienced are increasingly relying on AI tools. Which means that the skill set that is increasingly becoming more important for the deployment of AI code is also slowly eroding. Luciano Noison, a lead developer, wrote a blog post warning other developers to be cautious about making AI a key part of the workflow. He recounted how he had started relying on the AI tools provided by his company because he had started to feel left behind. About a year later, he removed those tools from his workflow entirely. Why? Because while he was working on a personal project, when he didn't have access to his company's provided tools, he realized that he had lost the instinct for tasks that used to be second nature for him. So this is what's at stake, the structural integrity of some of the most popular platforms we all use. And also, the skill sets of the architects who prop it all up. It's an entirely solvable problem because you see, it's not like these developers are anti AI. Even the 1500 Amazon petitioners were petitioning for Claude code as an alternative. The problem lies in the fact that companies are prioritizing adoption to present nice looking numbers to share holders and partners and not the needs of their own engineers. So what is needed is a balance between quantity and quality and between mediocrity delivered quickly and real skill developed over time. What you're listening to is just a small sample of our subscriber only offerings. A full subscription offers daily long form feature stories, newsletters and a whole bunch of premium podcasts. To subscribe, head to thecand.com and click on the red subscribe button on the top of the Ken website. Today's episode was hosted and produced by my colleague Rachel Varghese and edited by Rajiv Sien. [Music] [Music]
Podcast Summary
Key Points:
A 2025 study by research non-profit meter found that developers using advanced AI coding tools were 19% slower than those without, contradicting their own and perceived productivity estimates.
Despite this, major companies like Citibank, SAP, and Google have widely adopted AI coding tools, citing productivity gains, while a later study by Faro's AI showed increased code volume and task closure but no meaningful company productivity gains due to higher bug rates and review bottlenecks.
Indian IT firms (Infosys, Wipro, TCS, Cognizant) are heavily deploying AI tools via partnerships with Microsoft, claiming 20-40% productivity gains, despite thinning evidence.
Companies like Shopify and Amazon have mandated AI tool usage, with Amazon requiring 80% weekly usage of its internal tool Kiro, leading to engineer frustration and petitions for alternatives.
AI-generated code has caused significant outages (e.g., AWS 13-hour interruption, Amazon e-commerce six-hour outage) due to errors and autonomous actions by AI agents.
The Faro study reveals that bugs per developer are up over 50%, production incident probability has tripled, median review time increased by over 400%, and over 30% of code goes live without human review.
There is a shortage of senior engineers qualified to review AI code, leading to "code overload" and tech debt, as junior developers rely on AI without building fundamental debugging skills.
The shift risks eroding the skill sets of experienced developers and the structural integrity of platforms, as companies prioritize adoption metrics over real productivity and quality.
Summary:
The transcription discusses the complex impact of AI coding tools on software development productivity. In early 2025, a study by meter found developers were 19% slower with AI tools, contradicting their expectations. Despite this, major companies globally adopted AI tools, citing productivity gains.
A later Faro's AI study of 22,000 developers confirmed faster individual output but no company-wide gains, as AI code introduced more bugs, tripled incident probability, and increased review time by over 400%. Companies like Amazon and Shopify mandated AI tool usage, leading to engineer frustration and significant outages due to AI errors. Indian IT firms also heavily deployed these tools, claiming high productivity gains.
The transcription highlights a growing "tech debt" where junior developers rely on AI for code generation but lack debugging skills, while senior engineers are overwhelmed by reviewing poor-quality code. This threatens long-term skill development and platform integrity. The core problem is companies prioritizing adoption metrics over real productivity and engineer needs, risking a future with too much code and too few experts to manage it.
FAQs
The study found that developers using AI tools were actually 19% slower, despite estimating they would be 24% faster and feeling 20% faster.
Faro's study of 22,000 developers across 4,000 teams found that while developers write more code and close more tasks, bugs per developer are up over 50%, review time increased over 400%, and the probability of production incidents more than tripled.
In December 2025, AWS suffered a 13-hour outage in China after an engineer allowed Amazon's AI tool Kiro to make system changes, which deleted and recreated an environment. Amazon denied the cause was AI-generated bad code.
The volume of code produced by AI has overwhelmed reviewers, leading to over 30% of code reaching production without human review. There is also a shortage of application security engineers qualified to monitor AI code for risks.
Engineers using GitHub Copilot at Flipkart generated code that would automatically empty customers' shopping carts after adding five items, which was caught by a senior programmer just before deployment.
Junior developers rely on AI from day one, so they become good at generating average code but poor at debugging, missing the ability to trace errors from first principles. This creates a future shortage of senior developers who can review and fix AI-generated code.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.