Codex – Devstyler.io https://devstyler.io News for developers from tech to lifestyle Thu, 02 Apr 2026 10:25:26 +0000 en-US hourly 1 https://wordpress.org/?v=6.8.5 SonarSource Bets the Future of AI Coding Needs More Than Generation https://devstyler.io/blog/2026/04/02/sonarsource-bets-the-future-of-ai-coding-needs-more-than-generation/ Thu, 02 Apr 2026 09:03:47 +0000 https://devstyler.io/?p=136528 ...]]> With three new open beta products built around what it calls the Agent Centric Development Cycle, SonarSource is trying to solve a growing problem in software development: AI can write code fast, but that does not mean the code is trustworthy.

SonarSource unveiled three new open beta products — Sonar Context Augmentation, SonarQube Agentic Analysis and SonarQube Remediation Agent — designed to help teams guide, verify and fix AI-generated code throughout the development cycle. The company’s message is clear: as coding agents produce more software at much higher volume, the next battleground will not be generation alone, but whether organizations can trust, control and maintain what those agents create. 

Why This Matters for Users

For users, the benefit is practical rather than theoretical. AI coding tools can already generate large amounts of code quickly, but Sonar argues that speed often comes with more issues, more complexity and more technical debt. Its new tools are meant to reduce that burden by giving agents better context before they write code, checking their work while they are generating it, and fixing issues automatically before developers have to spend time cleaning them up by hand. 

The Competitive Difference: Sonar Is Selling a Control Layer

That is what separates SonarSource from many competitors chasing the AI coding boom. Plenty of vendors focus on helping agents generate code faster. Sonar is focused on what happens after that moment — whether the output aligns with architecture, passes quality and security checks, and can be repaired systematically without dragging down engineering teams. In other words, Sonar is not trying to be the coding agent itself. It wants to be the trust and verification layer around agent-driven development. 

The AC/DC Framework Behind the Launch

Sonar is packaging the launch around what it calls the Agent Centric Development Cycle, or AC/DC, a four-stage framework for AI-generated software: Guide, Generate, Verify and Solve. The idea is that AI agents should not operate as black boxes. They should first receive project-specific rules and architectural constraints, then generate code in a sandboxed flow, then have that code verified through deterministic analysis, and finally feed identified issues into a repair loop. That cycle, Sonar argues, is what turns AI coding from a novelty into an enterprise-ready process. 

Context Before Code

The first new product, Sonar Context Augmentation, is aimed at one of the most common weaknesses in AI coding: agents often lack awareness of the standards, structures and boundaries of the codebase they are working in. Sonar says the product injects relevant, real-time project context from SonarQube directly into the agent workflow, so the model understands what rules apply before it writes code. For customers, the value is not just cleaner output. Sonar says early benchmarks showed better build pass rates, better test pass rates, less code duplication, lower cognitive complexity and fewer tool calls and tokens, which could also mean lower operating costs. 

Catching Problems Earlier

The second product, SonarQube Agentic Analysis, moves code analysis into the agent’s generation loop instead of waiting for a failed pull request or human review. That could be meaningful for users because it shifts error detection upstream. If the code introduces a security risk, logic flaw or maintainability issue, the agent can see it and correct it in real time. The promise is that developers spend less time acting as cleanup crews for AI mistakes and more time on architecture and higher-value work. 

Fixing Technical Debt at Scale

The third product, SonarQube Remediation Agent, takes aim at both new issues and old backlog problems. For fresh pull requests, it can generate fixes as soon as SonarQube flags an issue. For older codebases, Sonar says it can work systematically through accumulated vulnerabilities, reliability issues and maintainability problems by opening one pull request per issue. That gives developers reviewed, ready-to-merge fixes without forcing automatic changes into production. The important distinction is that Sonar says every generated fix is re-scanned by its analysis engine before it reaches the developer, which strengthens its position as a verification-first platform rather than a blind automation tool. 

A Timely Message as AI Code Quality Comes Under Scrutiny

Sonar is also leaning on research to support its case. In the post, the company cites peer-reviewed Carnegie Mellon research covering 807 open-source projects that had adopted Cursor. Sonar says the study found a temporary productivity boost from agent usage, but by the third month that boost had faded, while code analysis warnings rose 30 percent and code complexity climbed 41 percent. For technology buyers, that is the core tension Sonar is trying to monetize: AI may increase output, but without stronger quality controls it can also increase long-term drag on development. 

Why Enterprises May Find This More Useful Than Another Coding Copilot

That framing could resonate especially with larger organizations that are already experimenting with Cursor, Claude Code, Codex, Gemini and GitHub Copilot but are concerned about compliance, maintainability and architectural drift. Sonar’s advantage is that it already has a long-standing position in code analysis and quality gates. Rather than asking customers to adopt yet another standalone AI coding product, it is extending that existing authority into the agentic era. For customers already using SonarQube, the transition may feel less like buying a brand-new category and more like upgrading an existing control point to meet AI-era demands. 

Image: Sonar 

]]>
OpenAI’s GPT-5.4 Targets Real Work: Build Apps Faster, Automate Tests, and Ship Business-Ready Docs With Agentic AI https://devstyler.io/blog/2026/03/06/openai-s-gpt-5-4-targets-real-work-build-apps-faster-automate-tests-and-ship-business-ready-docs-with-agentic-ai/ Fri, 06 Mar 2026 13:23:58 +0000 https://devstyler.io/?p=135015 ...]]> OpenAI has released GPT-5.4, a new flagship model it says is optimized for agentic workflows, combining stronger reasoning and coding with native computer-use capabilities—the ability to operate software via screenshots and mouse/keyboard actions—alongside support for up to 1 million tokens of context for long-horizon tasks.

A model built for “do the work” agents

In its announcement, OpenAI frames GPT-5.4 as its first general-purpose model shipping with state-of-the-art computer use, aimed at developers building agents that can complete real tasks across websites and software systems. The company highlights use cases such as automating workflows across apps, and notes the model can drive computer interactions directly and also write automation code via tools like Playwright.

OpenAI also says GPT-5.4 is more token efficient than GPT-5.2—using fewer tokens to solve problems—positioning it as both faster and cheaper in practice for certain workloads despite higher per-token pricing.

Benchmark claims emphasize professional work, tools, and desktop navigation

OpenAI’s post spotlights gains across a mix of “knowledge work” and agent benchmarks. It reports 83.0% “wins or ties” on GDPval (a professional work eval spanning 44 occupations), compared with 70.9% for GPT-5.2.

For computer-use tasks, OpenAI reports 75.0% success on OSWorld-Verified, up from 47.3% for GPT-5.2, and notes this exceeds 72.4% human performance in the benchmark notes.

Mercor: “Top of the leaderboard” for professional services agent work

OpenAI’s announcement includes early customer validation from Mercor CEO Brendan Foody:

GPT-5.4 is the best model we’ve ever tried. It’s now top of the leaderboard on our APEX-Agents benchmark, which measures model performance for professional services work. It excels at creating long-horizon deliverables such as slide decks, financial models, and legal analysis, delivering top performance while running faster and at a lower cost than competitive frontier models.

— Brendan Foody, CEO at Mercor

OpenAI also claims improved web-browsing and tool-use performance, including higher results on BrowseComp and Toolathlon, as part of its pitch that GPT-5.4 is better at selecting and operating tools in complex workflows.

What’s in it for developers

For software teams, OpenAI is positioning GPT-5.4 as a more capable “agentic” engine for end-to-end engineering work—especially when paired with tooling and computer-use interfaces.

GPT-5.4 is designed to take on longer development loops without losing context, thanks to its up to 1M-token window, which can help when the relevant code spans large repositories, extensive logs, multi-step incident timelines, or large test outputs.

OpenAI also emphasizes improvements in coding and debugging, with GPT-5.4 used inside Codex and integrated into workflows where the model can not only propose code changes but also drive tooling through a computer-use layer—opening the door to agents that can run commands, inspect outputs, and iterate.

For QA and test engineering, the company’s “computer use” capability is a notable shift: GPT-5.4 can be used to generate automated UI testing flows (for example, by producing Playwright-style scripts) and to execute multi-step test procedures where results must be validated and corrected across iterations. OpenAI’s OSWorld-Verified results are presented as evidence the model can reliably operate desktop environments for task completion.

OpenAI also claims GPT-5.4 is “more token efficient” than GPT-5.2, which can matter for developer workloads where tools generate verbose outputs (stack traces, logs, diffs) and cost is linked directly to tokens processed.

What’s in it for business teams

OpenAI is also aiming GPT-5.4 squarely at knowledge work—particularly tasks that combine research, synthesis, and output formatting into business-ready deliverables.

The company says GPT-5.4 improves generation and editing of documents, spreadsheets, and presentations, and ties those improvements to its “long-horizon” planning approach in ChatGPT through GPT-5.4 Thinking, which can surface an upfront plan for complex tasks that users can steer.

In practical terms, that’s the workflow Mercor describes: producing slide decks, financial models, and legal analysis as complete, multi-step deliverables, rather than short answers.

OpenAI’s own metrics are meant to reinforce the “business usefulness” angle. The company reports 83.0% wins-or-ties on GDPval, designed to measure task performance across dozens of occupations, compared with 70.9% for GPT-5.2.

The model’s computer-use capability also matters outside engineering: GPT-5.4-powered agents could navigate web dashboards, move data between tools, generate reports, and update systems of record—work that’s often manual across operations, finance, HR, and customer support. OpenAI frames this as part of its broader shift toward agents that can “do” work across applications, not just answer questions.

Reliability and rollout

On accuracy, OpenAI says that—based on de-identified prompts where users flagged factual errors—GPT-5.4’s individual claims are 33% less likely to be false and its full responses are 18% less likely to contain any errors, relative to GPT-5.2.

OpenAI says GPT-5.4 Thinking is rolling out in ChatGPT to Plus, Team, and Pro users, replacing GPT-5.2 Thinking, which is scheduled to be retired on June 5, 2026 after a three-month legacy period.

Image: OpenAI

]]>
OpenAI Hires an Army of Contractors – What Will Happen to Programming? https://devstyler.io/blog/2023/02/02/openai-hires-an-army-of-contractors-what-will-happen-to-programming/ Thu, 02 Feb 2023 08:27:42 +0000 https://devstyler.io/?p=99982 ...]]> OpenAI has increased hiring around the world, bringing on about 1,000 remote contractors in regions such as Latin America and Eastern Europe in the past six months, according to insiders, Semafor reports.

The article states that about 40% of the positions the company is seeking are computer programmers who create data for OpenAI models to learn software engineering tasks.

This news comes about a week after Microsoft announced it was laying off 10,000 jobs as well as an incredible multi-billion dollar investment in OpenAI, the company that created ChatGPT.

It would also be good to recall that OpenAI released a tool called Codex in August 2021, designed to translate natural language into code.

“They most likely want to feed this model with a very specific kind of training data, where the person provides a step-by-step layout of their thought process,”

a developer told Semafor, who was on a five-hour unpaid programming test for OpenAI.

He has asked to remain anonymous so as not to jeopardize future job opportunities.

OpenAI appears to be creating a dataset that includes not only lines of code, but also the human explanations behind them, written in natural language. The developer in question was asked to tackle a two-part series of tasks.

First, he was given a coding problem and asked to explain in written English how he would approach it. The developer was then asked to provide a solution. If he found a bug, OpenAI told him to describe in detail what the problem was and how it should be fixed, rather than just fixing it.

]]>
ChatGPT Can Detect and Fix Source Code Bugs https://devstyler.io/blog/2023/01/30/chatgpt-can-detect-and-fix-source-code-bugs/ Mon, 30 Jan 2023 08:50:34 +0000 https://devstyler.io/?p=99648 ...]]> ChatGPT detects and debugs source code using standard machine learning approaches, reports Analytics Insight.

Its main advantage over other AI methods and models is its unique ability to converse with humans, allowing it to improve the correctness of answers.

“We find that ChatGPT’s bug fixing performance is competitive to the common deep learning approaches CoCoNut and Codex, and significantly better than the results reported for the standard programme repair approaches,”

the researchers write in a new arXiv paper, which New Scientist first spotted.

Although ChatGPT’s ability to solve coding problems is not new, the researchers emphasize that its unique ability to talk to people gives it a potential advantage over other approaches and models.

According to Meta’s Head of Artificial Intelligence Jan Lekun, ChatGPT is built on the Transformer architecture, which was developed by Google.

In the code debugging examples, OpenAI highlights ChatGPT’s dialog capability, where it can ask for clarifications and get hints from the human to arrive at a better answer. Reinforcement learning with human feedback (RLHF) was used to train the large language models that power ChatGPT (GPT-3 and GPT 3.5).

ChatGPT’s ability to discuss may help it arrive at a more correct answer, the researchers note that the quality of its suggestions is not yet known. Therefore, they want to evaluate ChatGPT’s debugging capabilities.

The implications for developers are not yet clear. ChatGPT-generated responses were recently banned from Stack Overflow due to their low-quality but plausible-sounding nature. The Wharton professor finds that ChatGPT can act as a “smart consultant” (one that produces elegant but often incorrect answers) and promote critical thinking in MBA students.

]]>
AI Tools That Can Generate Code to Help Programmers https://devstyler.io/blog/2023/01/03/ai-tools-that-can-generate-code-to-help-programmers/ Tue, 03 Jan 2023 08:24:24 +0000 https://devstyler.io/?p=97439 ...]]> We recently shared with you an interview with GitHub CEO Thomas Dohmke for Computer Weekly, in which he argued that AI will not replace developers, no matter how quickly it gains popularity in IT circles. And contrary to the plethora of speculation spreading across the Internet and beyond, Dohmke believes that AI would accelerate productivity, but it wouldn’t entirely replace the need for humans.

And because AI can be a powerful weapon in the hands of developers if they know how to use it properly, today we’ve chosen to introduce you to some of the best AI tools, according to Marktech Post, that are currently available to developers.

AI Tools That Can Generate Code to Help Programmers

OpenAI Codex

GitHub Copilot, a tool from GitHub to produce code inside common development environments such as Neovim, VS Code, JetBrains, and even in the cloud with GitHub Codespaces, is powered by OpenAI Codex, a model based on GPT-3. It claims it can write code in at least 12 different languages, including BASH, JavaScript, Go, Perl, PHP, Ruby, Swift, and TypeScript. The algorithm is trained on trillions of lines of publicly accessible code from places like GitHub repositories.

Tabnine

Although Tabnine is not an end-to-end code generator, it amps up the integrated development environment’s (IDE) auto-completion capability. Jacob Jackson created Tabnine in Rust when he was a student at the University of Waterloo, and it has now grown into a complete AI-based code completion tool.

CodeT5

Researchers at SalesForce created the open-source programming language paradigm known as CodeT5. The T5 (Text-to-Text Transfer Transformer) framework from Google is its foundation. The researchers used approximately 8.35 million instances of code, together with user comments, from openly available GitHub projects to train CodeT5. The bulk of these datasets was obtained from the CodeSearchNet dataset, containing two C and C# datasets from BigQuery, along with Ruby, JavaScript, Go, Python, PHP, and C and C#.

Polycoder

OpenAI’s Codex has a competition in the form of a Polycoder. The model, created by scientists at Carnegie Mellon University, is based on OpenAI’s GPT-2, which was trained using a 249 GB codebase developed in 12 different programming languages. The creators of PolyCoder claim that the software can write C more precisely than any other model, including Codex. Polycoder is one of the earliest open-source code-generating models, even if most code generators are not.

Cogram

Cogram is a startup from Berlin’s Y-Combinator incubator that creates code for data scientists and Python programmers using Jupyter Notebooks and SQL queries. English-language queries may be written by data scientists and converted by the tool into sophisticated SQL queries with joins and grouping. It works with MySQL, SQLite, PostgreSQL, and Amazon Redshift.

GitHub Copilot

An AI tool called GitHub Copilot may assist you in producing better code. It can create code for you and aid in your comprehension of other people’s code. GPT-3 and OpenAI Codex power Github Copilot.

DeepCode

DeepCode is a code review tool powered by AI that examines your code and makes suggestions for improving it. Code completion, refactoring, and lining are among its many capabilities. For open-source projects, DeepCode is free, while a premium membership is available for private enterprises.

Kite

For Python, Kite is a free AI-powered code completion tool. You will get real-time code completions thanks to machine learning. For a fee, Kite also provides access to premium services, including sophisticated code analysis and refactoring tools. Kite stands out from the competitors since it supports more than 16 languages and 16 code editors. The regular updates to Kite make this machine-learning code aid more dependable and economical than the competition.

TabNine

An AI-powered code completion application called TabNine employs deep learning to offer possible code completions. This is accomplished by taking a piece of code and then offering comparable bits of code that may be used for the same issue. In addition to supporting more than 50 programming languages, TabNine is free.

CodeWP

The WordPress code generator CodeWP was created by Isotropic, which is who we are. This platform offers JS and PHP support and settings tailored to well-known plugins like WooCommerce and major page builders. It is particularly designed and optimized for those who construct WordPress websites.

Wing Pro

This editor examines static and dynamic code to provide excellent, context-relevant recommendations. Additionally, it offers you a better editing experience with a clever error-checking tool. The editor’s auto-completion functionality and built-in Python shells are both available. This tool includes a Source Assistant that constantly updates to provide inline documentation, type information, and call suggestions. As you code, it also automatically inputs function and method parameters. Wing Pro also allows you to browse the invocation and appropriately put your parameters.

]]>