Hype Cycle for Platform Engineering, 2026

14 May 2026 - ID G00846917 - 153 min read
By Cary Pillers, Bill Blosen,  and 1 more
Platform engineering focuses on building and operating a self-service internal developer portal and platform to enhance developer experience and productivity. Software engineering leaders should use this research to identify, plan and implement cost-optimized, AI-augmented practices and technologies.

Analysis


What You Need to Know

Platform engineering is a critical foundation for adopting, scaling and governing modern development practices, including agentic AI. Use this research to understand which technologies to integrate into the platform and scale to the organization.
As AI agents proliferate, platform engineering is essential to scale developer productivity and business value while embedding governance and cost by default.
According to the Gartner Software Engineering Survey for 2026, 81% of respondents consider platform engineering to drive moderate to high value in automating security and compliance workflows.1 Platform engineering simplifies the creation of secure and compliant software by embedding necessary tools and practices that developers and AI agents always follow by default.
Manage the rapid scaling of AI by standardizing tools, automating governance, enabling safe adoption and empowering developers and AI agents to operate efficiently, responsibly and autonomously. This enables faster delivery of business value while being able to manage the risks and costs of using AI.

The Hype Cycle

Over the last year, Gartner has seen platform engineering reach widespread adoption in the mainstream market and begin broad uptake in the late majority. Platforms increase standard adoption, speed up time to market, and advance teams toward the target state architecture. According to the Gartner Software Engineering Content Survey for 2026, AI strategy and developer productivity dominate engineering organizations’ priorities.1 AI-native development is demanding a shift in how platform teams meet the needs of their users (developers and AI agents) and deliver value toward AI goals. AI agents are increasing the speed of software development, leading to an increase in developer productivity. Platform engineering provides the standards and automations to manage AI at scale, freeing developers to deliver value to the business.
The innovations in this Hype Cycle fall under five themes:
  • AI-native development: The advent of AI-native development has added importance to platforms and platform engineering as a key foundation to rapid adoption of AI. According to Top Strategic Trends in Software Engineering for 2026, platform teams are a prerequisite for taking full advantage of this trend.
  • Developer productivity: Enhancing developer productivity remains a top priority for leaders. Platform engineering enables developer productivity by abstracting underlying complexity. Over the last year, Gartner has seen platform engineering reach widespread adoption in the mainstream market and begin broad uptake in the late majority.
  • Compliance by default: This starts with building software that is secure and compliant by default. Platform engineering simplifies the creation of secure and compliant software by embedding necessary tools and practices that developers and AI agents always follow by default.
  • Navigating the complexity of DevOps and cloud-native architectures: The continuous delivery of microservices using containers, service meshes and Kubernetes increases environment sprawl, making it harder to track service ownership and manage infrastructure at scale. As they use AI to write software, developers and citizen developers need platforms that manage this complexity and at scale.
  • Cost management: This involves using platform engineering to make costs visible and making the management of costs a requirement. It also involves enabling proactive management of AI and cloud costs by integrating cost control mechanisms directly into the paved roads.
Figure 1: Hype Cycle for Platform Engineering, 2026
Hype Cycle for Platform Engineering, 2026, plots 39 innovations from the Innovation Trigger through the Slope of Enlightenment. Innovations range from agent experience to Model Context Protocol to platform engineering.

The Priority Matrix

The following transformational innovations will have a significant impact on an organization’s business models, driving new strategies and tactics. Teams deploying platform engineering should note the following:
  • AI-native development enables developers, and non-developers to use AI agents, asynchronously and autonomously to deliver both productivity improvements and a creativity boost.
  • Agent experience (AX) is a design and engineering discipline focused on preparing back-end systems to attract and serve AI agents. AX ensures APIs, data, documentation, workflows and interoperability standards are machine-readable, discoverable and reliable.
  • AI agent management platforms (AMPs) provide a unified interface to secure, monitor, and govern agents independent of where they are deployed.
Recommended actions include:
  • Implementing internal developer platforms to drive standards, automations and developer experience.
  • Using paved roads to provide best practices for developers to follow by default.
  • Using the product-centric model to drive changes to the platform that developers want and need.

Priority Matrix for Platform Engineering, 2026

BenefitYears to Mainstream Adoption
Less Than 2 Years2 to 5 Years5 to 10 YearsMore Than 10 Years
Transformational
High
Moderate
Low
Source: Gartner (May 2026)

Off the Hype Cycle

  • DevOps was removed as it fully matured
  • Microservices was removed as it has fully matured
  • Cloud Native Architecture was removed as it has fully matured
  • Innersource was removed due to low adoption
  • Green Software Engineering was removed as AI-driven energy demand now dominates the sustainability conversation. Leading companies like Google and Microsoft are focused on securing additional power to support AI capabilities, making incremental green software improvements insignificant relative to the scale of AI’s energy requirements.

On the Rise

Agent Experience

Analysis By: Brent Stewart, Paige Kirk, Sarah Baumunk
Benefit Rating: High
Market Penetration: 1% to 5% of target audience
Maturity: Emerging
Definition:
Agent experience (AX) is a design and engineering discipline focused on preparing back-end systems to attract and serve AI agents. AX ensures APIs, data, documentation, workflows and interoperability standards are machine-readable, discoverable and reliable. Like UX for humans, AX makes software understandable, portable and actionable for agents through standardized interfaces, efficient execution patterns and governable contracts across systems.
Why This Is Important
Poor AX degrades human UX: when agents can’t reliably complete tasks, people get inconsistent outcomes and switch providers by redirecting their agents elsewhere. Also, as agent autonomy increases, agents will “shop” for systems that maximize task success (and avoid operational surprises), making AX a competitive differentiator in what agents choose to use.
Business Impact
Better AX makes your system easier for agents to choose, trust, and use. Clear interfaces based on AI-focused standards such as MCP, machine-readable intent, and predictable workflows reduce retries, failures, and compute waste so agents complete more tasks per dollar. The result is more agent-initiated calls, higher conversion and throughput, and new revenue via API consumption, marketplace distribution, and partner integrations.
Drivers
  • Protocols are defining the agentic web. Open standards like MCP are making systems easily pluggable into agents via consistent tool/data connections, which raises the premium on agent-friendly interfaces, schemas, and workflows.
  • Interoperability is becoming a competitive imperative. The Linux Foundation-hosted Agentic AI Foundation and backing from OpenAI, Anthropic, and Block signal a market push toward shared agent standards — making “agent choice” and experience a first-order concern.
  • Cost, reliability, and governance constraints are forcing structure. As agents scale, organizations must optimize for goal achievement with robust recovery from internal/external failures, while keeping humans central to enforce normative, financial, and agency guardrails. Governance must cover not just observable workflows, but agent security/access controls and drift detection. FinOps pressure from wide per-run cost variance further drives demand for predictable, measurable, and auditable execution.
  • Platform vendors are shipping agent-native runtimes and interoperability layers, accelerating the shift from chat UIs to tool-driven execution with durable state (memory). In this model, AX becomes a deciding factor in whether agents can discover capabilities, execute tasks efficiently, and recover predictably when conditions change.
  • Agent usage diverges from human usage. As agents expand in scope and capability, they stop behaving like “fast users” of human UIs and instead execute goal-driven, multistep workflows with emergent behaviors (tool chaining, retries, workarounds). This requires teams to optimize systems for agent-native interaction patterns and guardrails, not just human-inspired workflows.
Obstacles
  • Fragmented standards and semantics: MCP/tool calling is converging, but schemas, permissions, error models vary, so agents can’t generalize reliably across vendors.
  • Reliability and observability are immature: Few teams have SLOs, traces, evals, etc., making failures hard to debug, expensive to harden, and risky to automate.
  • Governance and safety overhead: Least-privileged access, auditability, data boundaries, and human override are non-negotiable in enterprises, but they add significant effort.
  • Legacy UX/API coupling: Systems optimized for humans require refactoring into explicit, machine-readable contracts, Often across many teams and years of tech debt.
  • Cost and value uncertainty: Agent runs can be materially more expensive than traditional automation due to model inference, tool retries, and orchestration overhead.
  • Cultural trust and privacy: Many users already defend against perceived surveillance (e.g., ad blockers) and may try to block agents from acting on their behalf.
User Recommendations
  • Treat agents as top users: Map agent personas and journeys alongside human users. Build an AX backlog and track success, retries, and cost-per-success.
  • Make intent + workflows reliable: Use schema-first tool contracts, normalize goals into an intent model, and design API-first, replayable state transitions with traces from intent to tool calls to outcome for auditability.
  • Design for human agent experience (HAX) now, autonomy later: Ship mixed-initiative patterns that enable human-AI partnership and progressively “thin” GUIs (for most systems) into supervision consoles and exception queues.
  • Operationalize governance: Pilot an agent integration layer, run controlled A2A pilots, and harden quarterly with policy-as-code, circuit breakers, and rollback paths based on observed failures, drift, and compute waste.
  • Instrument AX like APM/RUM: Implement real and synthetic agent run monitoring with end-to-end traces, replayable state, and outcome/cost metrics.
Gartner Recommended Reading

FinOps for Agentic AI

Analysis By: Ashish Banerjee, Andrei Razvan Sachelarescu
Benefit Rating: Moderate
Market Penetration: 1% to 5% of target audience
Maturity: Emerging
Definition:
FinOps for agentic AI is the financial discipline of managing the volatile costs inherent to agentic AI. Unlike clear and predictable costs of traditional software, agentic AI triggers unpredictable variable expenses through many LLM calls, longer reasoning traces, larger contexts, tool retries and multiagent loops. Without rigorous financial guardrails, attribution and observability, these systems can spiral into unpredictable token spend and API charges with little insight into actual ROI.
Why This Is Important
FinOps for agentic AI is important as the costs of widespread AI agent use becomes highly variable due to context size, tool calls, retried prompts and agent swarms, not just traffic. It can help enforce budgets, model routing, caching, loop/timeout breakers and step attribution to forecast unit economics, cap tail risk and enable showback/chargeback with governance. It connects the cost of operations of AI agents to the business outcomes, such as worker productivity or developer efficiency.
Business Impact
Agentic AI changes the spend pattern from a predictable “per user per session” to an unpredictable “per decision path” that can branch, retry, call tools, expand context or involve multiple agents. FinOps for agentic AI provides cost-related policies, governance and correlations to business outcomes for AI agents. Providing outcome-connected cost intelligence enables AI leaders to make more informed decisions around AI and agent investments, such as prioritizing the highest ROI use cases.
Drivers
  • Explosive enterprise rollout of task-specific agents (forecast to jump to about 40% of enterprise apps by 2026 YE) increases autonomous LLM/tool calls and makes spending harder to predict (see Gartner Predicts 40% of Enterprise Apps Will Feature Task-Specific AI Agents by 2026, Up from Less Than 5% in 2025).
  • The agent-to-agent economy: By 2028, 60% of agents will autonomously interact with external systems, creating a complex web of cross-vendor costs (see AI Vendor Race: Multiagent Systems Will Force an Evolution in Hybrid and Multicloud FinOps Practices).
  • Proliferation of multiagent orchestration in platforms (agents delegating to other agents) multiplies steps, context, retries — creating “tail-risk” bills.
  • Reasoning-model adoption for complex planning or decision making (models trained to “think longer and harder”) raises token/compute intensity per task.
  • GPU and AI infrastructure supply dynamics drive high acquisition and operating costs, pushing FinOps teams to apply tighter governance and forecasting.
  • Democratized decision making: AI spend is shifting from IT to business units, who now fund nearly 30% of AI initiatives (see AI Portfolio: How to Vet, Prioritize and Fund AI Use Cases), requiring clear showback/chargeback models to maintain accountability.
  • Scale overload: Past “everyday AI,” agent and tool sprawl makes spend and entitlements opaque; a unified FinOps and ITAM practice creates a single inventory of agents/models/tools/licenses, assigns owners, enables showback/chargeback, and governs renewals/compliance, so complexity does not outpace value.
  • Maturing cost-control primitives (capacity reservations/PTUs, model routing and integrated observability) let FinOps actively balance quality, latency and cost at runtime — turning FinOps from reporting into a closed-loop control plane for agent workloads.
  • The availability of AI agent governance tools (gateways, catalogs, policy controls) emboldens businesses to scale quickly despite concerns around cost governance.
Obstacles
  • Rapid model/pricing churn: Frequently changing model tiers, token accounting and multivendor discount structures make forecasting and unit economics unstable.
  • Immature tooling: Limited routers, caching, budget enforcement and automated policy tests in many runtimes.
  • Sparse, inconsistent telemetry: Token/reasoning/tool costs often lack unified tags and per-step attribution across stacks.
  • Optimization trade-offs: Cost controls can degrade accuracy/latency, slowing adoption until patterns mature.
  • Data/privacy constraints: Logging prompts/tool outputs for cost/quality auditing can conflict with compliance.
  • Cross-team ownership gaps: App teams, platform, FinOps and security lack shared KPIs and operating model.
  • Talent/time scarcity: The people best positioned to add cost guardrails (agent builders) are also the scarcest; their time is prioritized for new capabilities, not cost instrumentation and optimization.
User Recommendations
  • Stand up agent cost telemetry: Log tokens, tool calls, latency, $ per step; enforce tagging by agent/workflow/environment.
  • Define unit KPIs: Cost per resolved task, success rate, agent value multiple (total value generated/agent cost), effective context utilization.
  • Implement guardrails: Max steps/tokens/$, timeouts, loop detection, fallback paths, human-in-the-loop for high risk.
  • Route smartly: Default to small models; invoke reasoning models only for hard decisions; cache, reuse results.
  • Control context: Summarize, externalize state, cap retrieved docs.
  • Build governance: Policy for tools/data access, approvals, audit trails; run automated evals before release.
  • Co-create cost-benefit models with IT finance that match real-time consumption with business results.
  • Operationalize agentic FinOps as a shared capability: Embed FinOps practitioners into the AI agent engineering team to run monitoring, cost-anomaly detection and guardrails — tied into accounting/showback, and not built in isolation.
Sample Vendors
Airia; Exostellar; Finout; Flexera; IBM
Gartner Recommended Reading

Infrastructure From Code

Analysis By: Hassan Ennaciri
Benefit Rating: Low
Market Penetration: 1% to 5% of target audience
Maturity: Embryonic
Definition:
Infrastructure from code (IfC) is an advanced automation approach that derives infrastructure needs directly from application code, thereby provisioning cloud resources based on policy. Developers’ code intent is interpreted — such as a database call or a message queue — to automatically provision secure, compliant resources. By shifting infrastructure “left” into the development workflow, IfC eliminates manual handoffs and ensures cloud environments are synced with application logic.
Why This Is Important
IfC merges infrastructure and application code into a common programming model and language. By autodiscovering requirements directly from business logic, IfC enables infrastructure to evolve with software. Developers focus on business logic and application architecture, while infrastructure code is autogenerated based on organizational standards. IfC abstractions aim to make it easier for application developers to adopt new cloud services using programmable software development kits.
Business Impact
IfC drives business impact by accelerating innovation velocity and the ability to experiment. By unifying application logic and infrastructure into one model, it allows teams to test new architectures and AI workloads in minutes rather than days. It minimizes operational risk by injecting security guardrails directly into code analysis, preventing costly breaches. This enables focus on feature development, transforming infrastructure from a bottleneck into a competitive engine for rapid growth.
Drivers
  • Improving innovation velocity: IfC removes the “waiting for infra” bottleneck. Because requirements are inferred from the app logic, teams can experiment with complex AI and data architectures in minutes, drastically shortening the time to market for new features.
  • Improving infrastructure agility: IfC builds on the foundations of platform engineering and SRE to further blur the boundaries between infrastructure and development. It reduces “cognitive load” by allowing developers to define what their code needs without requiring them to be experts in specific cloud provider consoles or syntax.
  • Platform engineering demand: The rise of platform engineering, site reliability engineering, and the focus on developer experience over the past decade have created a foundation for applying software engineering practices and patterns to infrastructure management. IfC builds on this foundation and aims to further blur the boundaries between infrastructure engineering and application development.
Obstacles
  • Technology immaturity: Most IfC tools are emerging startups lacking the feature completeness, stability, and massive community ecosystems of legacy IaC.
  • Variable cloud support: Abstraction layers require vendors to track cloud changes perfectly. This creates gaps where niche or new cloud services aren’t yet supported.
  • Coupling & governance friction: Merging app and infra code can bypass traditional audit workflows. In enterprises with manual policies, this tight coupling is often perceived as a security risk.
  • Technical debt: Autogenerated infrastructure creates transparency gaps. When the abstraction fails, SREs struggle to troubleshoot the unfamiliar underlying resources.
  • Upskilling hurdle: Moving to IfC requires teams to master entirely new programming constructs and SDKs, creating a steep learning curve for those used to declarative configuration.
User Recommendations
  • Experiment with multiple IfC tools before choosing one for production, due to the technology’s early maturity stage.
  • Evaluate tools based on their support for your cloud provider, services, programming language, use cases and regional vendor support.
  • Decide on long-term ownership of infrastructure assets by platform engineering teams or dedicated ops teams before using IfC in production.
  • Be cautious of relying solely on open-source offerings, as the market’s immaturity means future mergers could impact tool availability.
  • Test autogenerated infrastructure code or autodeployed infrastructure much like application code before production deployments. Engineers tend to assume that generated code is more accurate.
Sample Vendors
Ampt; Encore; Modal; Nitric; Shuttle; StackGen
Gartner Recommended Reading

Spec-Driven Development

Analysis By: Akis Sklavounakis
Benefit Rating: High
Market Penetration: Less than 1% of target audience
Maturity: Emerging
Definition:
Spec‑driven development (SDD) is a software engineering methodology in which semistructured, human-readable and machine‑interpretable specifications serve as the primary input of feature requirements in an AI agent-driven software development life cycle. SDD shifts the primary system of record to explicit intent. It operates through a structured life cycle that typically includes specification, planning and implementation phases.
Why This Is Important
AI is transforming software delivery by shifting from human-led workflows to human-supervised AI-driven delivery. AI tools increase developer velocity but introduce significant organizational risks. Generating code through loose prompts leads to inconsistent patterns, security vulnerabilities and unmaintainable codebases. SDD improves governance and enforces traceability and auditability by steering AI output with business intent and architectural standards through rigorous context management.
Business Impact
Effective SDD helps software engineers to:
  • Enforce architectural guardrails, security and compliance policies by design rather than treating them as an afterthought.
  • Preserve institutional knowledge during personnel transitions through defining specifications as the system of record for both humans and AI.
  • Decrease cycle time while preventing architectural drift, quality degradation and unmanaged risk.
  • Reduce token usage by presenting concise requirements to the LLM’s context window.
Drivers
  • As AI coding agents compress development time, the long-standing constraint of requirements ambiguity increasingly undermines delivery throughput. This realization forces a structural change toward treating specifications as a critical source of truth.
  • Enterprises demand predictable quality from AI code-generation tools.
  • Expanded context windows allow LLMs to process a variety of contexts, such as comprehensive architectural documents, effectively.
  • Organizations in regulated industries require strict traceability linking generated code directly back to original business requirements.
Obstacles
  • Tooling fragmentation: Tool-specific formats risk vendor lock-in.
  • Organizational debt: Enterprise workflows assume human operators and human-readable interfaces. Legacy systems lack the documentation needed to build initial specifications.
  • Maintenance overhead: Writing exact specifications requires high upfront effort. Teams risk spec drift if ad hoc changes are made directly to the code.
  • Skills gaps: Specification and context engineering skills are emerging and are hard to find and train.
  • Change resistance: The fundamental shift to AI-native SDLC can cause change friction, where engineers actively resist process changes that force them to write text files instead of code.
User Recommendations
  • Prevent architecture technical debt by mandating a spec-first approach for all AI-generated production code. Treat specifications as a version‑controlled source of truth that governs how code is generated, validated and evolved.
  • Enforce non-negotiable security, compliance, data privacy and architectural constraints by curating persistent context in the form of files across AI agents and related repo folders. This creates a governance layer that preemptively rejects noncompliant patterns, hardening security and architectural consistency by design, rather than relying on agents making assumptions to fill context gaps. Good context engineering practices apply to SDD as well.
  • Gradually scale the adoption of SDD by piloting workflows on greenfield initiatives before attempting complex modernization initiatives. Deploy SDD initially on small modules or application surfaces to refine the specification-to-code life cycle without the cognitive load of legacy technical debt.
Sample Vendors
AWS; Cognition (Windsurf); Cursor; For Good AI Inc. (Zencoder); GitHub; Tessl AI
Gartner Recommended Reading

Agent Management Platform

Analysis By: Tom Coshow
Benefit Rating: High
Market Penetration: 1% to 5% of target audience
Maturity: Emerging
Definition:
AI agent management platforms (AMPs) provide a unified interface to secure, monitor, and govern agents independent of where they are deployed. An AMP should also include the management of agent acquisition from marketplaces, libraries, and analytics regarding cost and performance.
Why This Is Important
As enterprises accelerate their adoption of AI agents and automation tools, the complexity of managing these resources is intensifying. AI agent management platforms are important because they bring structure, visibility, and control. The proliferation of AI agent development platforms drives organizations to deploy AI agents across multiple platforms and environments. This diversification brings innovation, but it introduces new challenges regarding integration, oversight, and risk management.
Business Impact
Enterprises deploying AMPs will have a significant advantage in preventing agent sprawl and instituting best practices in agentic AI. AMPs can:
  • Provide accountability to ensure alignment of investments to business outcomes.
  • Control the security relationship between agents, humans, and data.
  • Manage context and data inputs that affect the reliability of agent decision making.
  • Observe and alert on agent behavior changes due to changes in the LLM.
Drivers
  • Controlling AI agent sprawl: Controlling AI agent sprawl is essential to prevent redundant deployments, reduce operational costs, and mitigate security risks by maintaining centralized visibility over all AI agents.
  • Managing agent marketplaces and libraries: Managing agent marketplaces and libraries ensures that only vetted, compliant, and secure agents are procured and deployed, reducing the risk of shadow IT and maintaining quality standards.
  • Tooling catalogues: Tooling catalogues provide a comprehensive inventory of approved tools and integrations, enabling organizations to standardize solutions and manage dependencies effectively.
  • Securing agent behavior: Securing agent behavior is critical to protecting sensitive data and business processes with robust controls and monitoring, ensuring agents act within defined policies.
  • Observing and governing AI agent behavior across platforms: Observing and governing AI agent behavior across multiple platforms allows organizations to enforce policies consistently, detect anomalies, and maintain compliance in hybrid and multicloud environments.
  • Financial reporting: Financial reporting delivers transparency into AI agent and tooling costs, supporting budgeting, cost optimization, and ROI analysis for informed decision making.
Obstacles
  • Building or buying an AMP that manages enterprise AI agent development and deployment is a nontrivial exercise. This task requires leveraging other AI observability, security, and governance tools.
  • The diversity of the agent platform ecosystem will challenge the effectiveness of integrations and communications.
  • Budgets for governance technologies lag the desire for a solution.
  • The market currently provides many different levels of AMP capabilities. Many of these solutions only manage agents on their platform.
  • The level of governance and observability across AI agent runtime platforms is handled differently by different providers and goes to different levels of detail and control.
  • Scalability and cost instability. Cost control, rate limiting, and workload optimization mechanisms are not yet fully mature.
  • Weak governance and guardrails. Enterprise-grade agent compliance and risk management are still evolving.
User Recommendations
  • Prioritize unified monitoring solutions as a key component and the starting point of AMP deployment.
  • Look for observability and governance tools with deep integrations across platforms that enable stopping an agent from taking a bad action and provide the ability to enforce guardrails.
  • Focus on business-critical agents and those that are exposed outside the governance controls of their host platform.
  • Project the needs of your organization three years out. How will you manage 500 AI agents spread across different departments, built on different platforms, and acquired from different marketplaces?
  • Use AMPs to provide enterprisewide financial reporting on AI agents, regardless of which platform they operate on. Provide cost and ROI dashboards at the agent detail and summary levels.
  • Define agent governance from Day 1. Define ownership, approval workflows, and monitoring standards upfront. Implement logging, audit trails, and fallback mechanisms.
Gartner Recommended Reading

Eval-Driven Development

Analysis By: Manjunath Bhat
Benefit Rating: High
Market Penetration: 5% to 20% of target audience
Maturity: Emerging
Definition:
Eval-driven development (EDD) is an approach to systemically build quality measures such as explainability, security, safety and trust into GenAI and agentic AI applications. EDD guides the design, development and testing of these applications via a set of task-specific evaluations (evals) that assess application outputs against predefined quality criteria. AI engineers can use evals to inform context engineering, model selection and design patterns and mitigate the risks of nondeterminism.
Why This Is Important
GenAI applications and agents are inherently nondeterministic and opaque. This makes it difficult to measure and improve their reliability and alignment using traditional testing methods. In addition, the vast range of possible inputs and outputs in these applications makes it impossible to create test cases for every scenario. EDD helps resolve these challenges by using a rubric to evaluate how well the application performs on relevant metrics (e.g., correctness, accuracy, fairness, cost).
Business Impact
EDD applies test-driven development (TDD) principles to GenAI by defining and executing evals with observability in development and production. This approach proactively identifies risks and mitigates reputational damage. EDD helps engage business stakeholders early in the development life cycle to define success metrics, replacing guesswork with data-driven quality engineering methods. The increasing autonomy of agentic applications makes EDD critical to their development and delivery.
Drivers
  • Building predictability into fundamentally unpredictable systems: GenAI models are inherently probabilistic, making predictability a challenge. To tackle this, AI application developers are using evals as “fitness functions” to ensure AI systems adhere to defined quality standards in the organization. EDD helps continuously identify and address risks due to data and model drift.
  • Need to comply with global AI regulations: The U.S. executive order on AI (EO 14179) calls for AI development and use without bias such as “ideological biases” or “engineered social agendas.” The NIST AI Risk Management Framework outlines transparency, explainability and fairness as essential characteristics of trustworthy AI. EDD enables organizations to comply with these regulatory requirements and standards by identifying and mitigating risk factors using automated techniques.
  • Limitations of traditional testing methods: Traditional testing methods that rely on deterministic expectations fail to address the variability inherent in large language model (LLM) outputs. This limitation creates a significant gap in quality assurance for GenAI applications. Without a structured testing approach specialized for GenAI applications, teams resort to subjective evaluations that primarily rely on human “gut instincts.”
  • AI agents, multiagent workflows and retrieval-augmented generation (RAG): Software engineering teams are increasingly tasked with building AI agents to address a wide range of business and operational needs. The 2025 Gartner AI in Software Engineering Survey revealed that 68% of respondents are either actively building or have already developed AI agents. However, ensuring the reliability of agentic workflows and RAG architecture in production is difficult as they can fail in unpredictable ways.
Obstacles
  • Need for “ground truth” datasets in select use cases: The scarcity of preexisting “ground truth” datasets creates a bottleneck for adopting EDD in select use cases (e.g., RAG) where we want to measure retrieval accuracy (e.g., Did AI retrieve the correct document?).
  • Lack of expertise: To build robust evaluations, software engineers must learn and become adept at AI engineering skills including metrics frameworks such as Ragas, G-Eval and GPT Estimation Metric-Based Assessment (GEMBA). These frameworks help quantify subjective measures of accuracy, consistency, fluency, faithfulness, coherence, relevance and precision.
  • Rapid pace of change: Developers must update their evaluation criteria when adopting new capabilities like multimodal models or agentic workflows. This necessitates experimenting and evolving their LLM evaluation rubrics and frameworks. Failure to do so makes EDD seem ineffective.
User Recommendations
  • Build developer expertise in EDD by budgeting for training courses and investing in evaluation tools for automating evaluation runs.
  • Use the 80/20 rule: Start with a small, cross-functional team (subject matter experts and engineers) to write down the system’s purpose in plain terms. Don’t wait for thousands of test cases. Start with 20 to 50 high-quality tasks drawn from real-world failures or manual tests.
  • Include evals as part of the “Definition of Done” for GenAI applications. Integrate evals into DevSecOps pipelines to serve as fitness functions and quality gates before production deployments.
  • Improve the quality of evals continually by using telemetry from production environments to refine evaluation metrics, techniques and datasets.
  • Define an “Acceptable Error Rate” per use case. A creative writing bot might have a 10% error budget, while a financial advisor bot might have a 0.1% budget.
Sample Vendors
Airia; Arize AI; Braintrust; Comet; Confident AI; CoreWeave; Deepchecks; Evidently AI; Fiddler AI; Galileo
Gartner Recommended Reading

GenAI Model Routers

Analysis By: Manjunath Bhat, Andrew Humphreys
Benefit Rating: High
Market Penetration: 5% to 20% of target audience
Maturity: Emerging
Definition:
Generative AI (GenAI) model routers are an intelligent middleware layer decoupling the interaction between AI applications and their model dependencies. They dynamically analyze and direct requests to the most appropriate model based on optimization criteria such as costs, latency, reliability and performance, optimizing for costs and quality. Model routers can be adopted as a stand-alone tool and/or included as a feature in AI application development platforms and AI gateways.
Why This Is Important
GenAI model routers optimize costs and response accuracy by directing requests to appropriate models. They ensure that each query is handled by models optimized for a specific need (e.g., creative writing, coding, image generation), which enhances output quality. GenAI model routers help achieve performance-cost trade-offs in harnessing model enhancements and innovations while limiting costs. They also help routing requests containing sensitive data to self-hosted, local models.
Business Impact
GenAI model routers can help optimize costs when deploying AI applications at scale by directing simpler queries to less expensive and smaller models. Smaller models can enhance an application’s responsiveness and reduce their resource usage. Model routing helps build AI applications that can adapt to evolving requirements by decoupling applications from underlying models and avoiding model lock-in.
Drivers
  • Optimizing costs: Routers can cut inferencing costs by diverting a subset of queries to smaller, more efficient models. This can be more cost effective than a single model strategy. For example, vLLM Semantic Router by Red Hat, an open-source router, improved accuracy by 10.2%, reduced latency by 47.1% and token usage by 48.5%, as measured on the MMLU-Pro benchmark.
  • Selecting the right model: The rapid pace of innovation, as evidenced by the versatility and volume of models, makes it challenging to choose the right models for a task. Static routing is both inefficient and ineffective due to varying model capabilities, drifting model behavior and token costs.
  • Building agentic AI workflows: Agentic AI applications can benefit from multiple models for different types of processing using the model router’s single abstracted model API.
  • Privacy, sovereignty and regulatory compliance: GenAI model routers enable directing prompts containing sensitive data to a different model, such as a potentially self-hosted or tenant-hosted model, while sending generic queries to a public API.
Obstacles
  • Multimodal inputs: GenAI model routers are still unproven for handling the complexity of multimodal inputs (e.g., audio, images and video). Currently, their ability to analyze prompts is only text based.
  • Potential latency implications for time-sensitive use cases: Model routers can introduce response delays while analyzing the input to make a routing decision in use cases, such as real-time event processing, that are extremely sensitive to response times.
  • Maturity on evaluations: Objectively proving that the optimal model is “good enough” for a specific use case requires a high degree of testing maturity (e.g., eval-driven development [EDD] practice). Teams struggling to build robust evals find it tough to keep up with the rapid release cycle of new models.
User Recommendations
  • Determine the need for GenAI model routers by estimating token usage for your agentic AI applications.
  • Optimize the cost of using frontier models by implementing model routers in conjunction with AI gateways to achieve optimal trade-offs between reliability, accuracy and cost.
  • Test router quality by using an EDD approach to ensure the responses are in line with expectations. Use evals to compare the performance and accuracy of model responses with and without routing. See Market Guide for AI Evaluation and Observability Platforms.
  • Adopt a tiered model strategy by assigning distinct model tiers for specific use cases — a small tier model for classification, data masking and basic formatting; a medium tier model for 70% to 80% of user interactions; and a large, reasoning tier model reserved for complex logic, math and code.
Sample Vendors
Airia; Amazon; Aurelio AI; LiteLLM; Lunar.dev; Microsoft; Not Diamond; Requesty; Traefik Labs; TrueFoundry
Gartner Recommended Reading

AI-Native Software Engineering

Analysis By: Manjunath Bhat, Mark Driver
Benefit Rating: Transformational
Market Penetration: 5% to 20% of target audience
Maturity: Emerging
Definition:
AI-native software engineering includes practices and principles optimized for using AI-native tools across the software development life cycle to accelerate software delivery. AI-native practices go beyond human augmentation and involve using AI agents for asynchronously and autonomously executing long-running tasks that span multiple use cases and diverse roles. AI-native ways of working can deliver both productivity improvements and a creativity boost.
Why This Is Important
AI-native software engineering practices enable teams to focus on meaningful work that requires critical thinking, creativity, and user empathy, rather than spending time on repetitive tasks. By adopting AI-native methods beyond coding tasks, software engineering leaders can maximize the impact of AI in the software development life cycle (SDLC) and realize greater return on their technology investments.
Business Impact
AI-native software engineering leads to maximizing the use of AI across the SDLC. The 2025 Gartner AI in Software Engineering Survey shows a striking contrast between teams maximizing AI use across the SDLC versus those minimally using it for fewer use cases. For example, 55% of respondents who use AI for 10 or more use cases see an increased rate of innovation while 53% cite an increase in user/customer satisfaction, and 61% report an increase in developer job satisfaction.
Drivers
The primary value drivers for AI-native software engineering include:
  • Need to go beyond productivity gains and use AI to drive innovation: As AI commoditizes code generation, the primary measure of engineering effectiveness is shifting from productivity to creativity and innovation. Used effectively, AI transforms the experience of people in upstream planning phases by serving as an ideation partner for roles such as product owners and user experience designers, enabling them to convert text prompts or visual sketches into prototypes and supporting better and faster decisions.
  • Emerging practices, such as spec-driven development and context engineering: Spec-driven development combined with agentic coding tools that have access to better quality context help software engineering teams get closer to the aspiration of implementing a zero-friction SDLC. AI agents use specs to guide planning and implementation, enabling multiple asynchronous workstreams to run in parallel and deliver faster cycle times.
  • Compounding the effects of AI-native development tools used in ensemble: AI-native development tools used in ensemble across the SDLC enable organizations to not only improve delivery speed but also build in quality guardrails. For example, AI code review tools, AI testing tools, AI code security assistants, and AI site reliability engineering tools continuously detect and remediate quality issues and incidents.
  • Need to elevate the human experience (for example, developer experience): AI tools can significantly enhance human experience by reducing cognitive load and facilitating “flow state.” By automating tedious tasks (such as writing unit tests or migration scripts) and minimizing context switching and information retrieval (such as “explain this error” or “find this dependency”), AI allows engineers to minimize distractions.
Obstacles
  • Blind trust in AI output: AI-native approaches create a new burden on developers and knowledge workers in general. Developers increasingly offload tasks to AI tools, which carry inherent risks of nondeterminism and hallucinations. Therefore, blindly trusting AI output without verification and explainability can potentially pose serious business risks, including reputational damage.
  • Increased security risk: AI tools expand the threat surface via MCP servers, agent skills, agent plug-ins, and IDE extensions, which increases the potential for unforeseen vulnerabilities and security breaches.
  • Developer burnout due to high-intensity work: While AI can free up time for creative work, there is a high risk that it actually intensifies work by drastically increasing baseline productivity expectations. Developers using AI-native techniques are more likely to experience increased cognitive load as they constantly validate AI output and take accountability for work they have not personally done.
User Recommendations
Adopt AI-native software engineering in three phases:
  • Phase 1: Resolve constraints by mapping the software delivery value stream. Identify systemic bottlenecks and resolve them by judiciously using AI where appropriate. Discover pain points and improve the experience for all roles, not just developers.
  • Phase 2: Reimagine the SDLC with asynchronous workflows. Parallelize tasks using asynchronous agentic workflows. Transform software delivery workflows — identify what steps can be eliminated and what stays the same. Enhance IDP capabilities to govern, monitor and control AI software engineering agents.
  • Phase 3: Realize zero-friction SDLC with autonomous software delivery. Implement autonomous self-correcting and self-improvement loops and address pitfalls by expanding platform support for autonomous delivery and operations. Implement appropriate human oversight for autonomous workflows based on business criticality, acceptable risk, and architectural complexity.
Sample Vendors
Amazon Web Services; Anthropic; Cognition; Cursor; GitHub; GitLab; Google; Harness; Lovable; OpenAI
Gartner Recommended Reading

Autonomous Workload Optimization

Analysis By: Pankaj Prasad, Hassan Ennaciri, Manjunath Bhat
Benefit Rating: High
Market Penetration: 5% to 20% of target audience
Maturity: Emerging
Definition:
Autonomous workload optimization (AWO) maximizes performance and efficiency by intelligently managing resources and configurations. It achieves this by scaling compute, network and storage assets; fine-tuning underlying runtime platform settings; and dynamically shifting workloads to help manage cost, while ensuring reliable performance. By replacing manual intervention with real-time, data-driven adjustments, it ensures an ideal balance between cost and performance.
Why This Is Important
Optimizing IT costs has always been on the head of I&O’s most critical priorities list. Additional pressures include ensuring optimal performance to consumers in an energy-efficient manner. AWO operations navigate complex trade-offs in real time, ensuring impactful results without sacrificing reliability or performance. By using algorithms to handle massive scale and competing priorities, organizations achieve sustainability and savings that are otherwise impossible to maintain manually.
Business Impact
Autonomous workload optimization tools enable organizations to optimize their IT costs and avoid wasting IT resources, while balancing the performance requirements of applications to ensure that consumers are not affected.
Drivers
Organizations adopt autonomous workload optimization in response to several priorities:
  • Cost optimization: Address the rising costs of running cloud-native architectures at scale and the need for efficient resource allocation without impacting performance. Autonomous workload optimization tools dynamically adjust resources to match demand, minimizing waste and maximizing cost-efficiency.
  • Performance optimization: Enterprises aim for an optimal range of performance in which cost is balanced, while avoiding a negative impact on consumer experience
  • Forecasting workload requirements: By mapping application resource utilization to demand patterns, especially in virtualized and cloud-native architectures, enterprises can dynamically project resource requirements, thereby optimizing their IT spending.
  • Efficiency goals: I&O leaders are committing to sustainability goals, and greater energy efficiency is one of their primary targets. This is also a driver for sustainable software engineering.
  • Algorithmic speed and scalability: Manual optimization is not sustainable in the long run, especially in environments requiring dynamic, large-scale optimization. The need for speed at scale is the driver for autonomous tools and processes.
  • Enhanced control via AI-driven interfaces: Agentic AI enables autonomous monitoring, reasoning and execution of fine-tuning tasks across complex systems, optimizing code and queries in real time. This provides organizations with the speed and ability to meet their performance targets.
Obstacles
  • Autonomy risks and trust barriers: Allowing autonomous systems to make real-time changes in production environments raises concerns about unintended consequences, compliance violations and loss of human oversight. Organizations hesitate to grant full “write” access without robust safeguards.
  • Lack of standardization: Many organizations rely on a mix of legacy systems, custom configurations and multiple cloud platforms, resulting in fragmented IT environments. This lack of standardization makes it difficult for AWO tools to access consistent data and communicate across systems, limiting optimization to basic actions.
  • Collaboration challenges: Data sharing among preproduction and production teams is crucial to avert negative impact to system stability from optimizations, especially since reliability expectations can be in conflict for independent versus integrated systems.
  • Disconnect with business stakeholders: Business leaders’ buy-in can be hampered when technical reporting doesn’t align with their preference for measuring operating performance. Resource optimization gains ideally need to be translated into dollar savings.
  • Limited observability: AWO relies on comprehensive, real-time visibility into workloads, dependencies and resource utilization. Gaps in monitoring or incomplete telemetry can prevent autonomous systems from making informed decisions.
  • Skills and change management: Adopting AWO demands new skill sets in AI operations, automation and governance.
User Recommendations
  • Begin by analyzing resource utilization against performance to identify patterns and correlations. Pilot autonomous optimization on noncritical workloads to prove value before wider implementation.
  • Use autonomous optimization initiatives to drive maturity in monitoring metrics and speed up standardization initiatives across IT architectures and processes to maximize the value of optimization objectives.
  • Improve collaboration across customer-facing, preproduction and production teams to ensure regular review and appropriate data exchange, including customer engagement, performance, IT resource requirement and utilization patterns.
  • Collaborate with business leaders to ensure that business and technical objectives are captured and relevant measurements and benefits are appropriately reported.
  • Embed autonomous workload optimization into the application deployment process to maintain ongoing visibility of outcomes, regularly review targets and continuously refine optimization goals.
Sample Vendors
Akamas; Avesha; Cast AI; Cisco (Opsani); CloudBolt (StormForge); DoiT (PerfectScale); Komodor; IBM; ScaleOps; Sedai
Gartner Recommended Reading

AI Gateways

Analysis By: Andrew Humphreys, Mark O'Neill
Benefit Rating: Moderate
Market Penetration: 1% to 5% of target audience
Maturity: Embryonic
Definition:
An AI gateway is a tool that acts as an intermediary between applications and various artificial intelligence services or models. Its purpose is to simplify and manage connectivity between AI applications, agents, LLMs and enterprise applications by providing a central point to enable security, governance and observability of AI workloads.
Why This Is Important
AI gateways manage interactions between AI applications, agents, AI models, and other applications by providing security, governance, observability and cost management. This reduces the risk of unexpected costs incurred from AI providers, and prevents private data in API traffic from being compromised or misused. An AI gateway may also enable an organization to manage usage of multiple AI models from different providers, rather than forcing reliance on one provider.
Business Impact
An AI gateways can help implement and manage prompt-based policy controls, track AI service use and costs, route across multiple LLMs, and implement request routing and support functions such as authentication and authorization, rate limiting, load balancing, caching, and logging for AI‑enabled interfaces, such as MCP Servers. These features enable organizations to have greater control over the interaction of AI applications and agents with existing applications and LLMs.
Drivers
  • Control and visibility of AI-consumable interfaces: Initial AI gateways focused on controlling access to LLMs. With the growth of adoption of AI integration protocols, AI gateways have expanded to provide governance of AI-consumable interfaces based on MCP and A2A.
  • Control access costs: Pricing models for AI services tend to be token-based, which presents a risk for businesses because the costs associated with use of these services can accrue rapidly, and requires special metering ability beyond an API gateway’s typical transaction-based metering. Organizations are looking to control access costs, and AI gateways can help by caching responses to limit duplicate calls and tracking and controlling access to the services.
  • Enable visibility of use of AI services: AI gateways provide a way to get greater visibility into the use of AI APIs across the organization by leveraging API management observability and analytics features for API usage.
  • Optimize access to AI engines: AI gateways can be configured to provide a single API in front of multiple different AI providers, such as multiple LLMs. This means that developers can access multiple AI services using the same API.
  • Enforce security on AI interactions: AI trust, risk and security management (AI TRiSM) ensures AI model governance, trustworthiness, fairness, reliability, robustness, efficacy and data protection. AI gateways address aspects of AI TRiSM, including model governance and reliability. In particular, protection of API keys issued by AI providers is vital. If an attacker gets access to an organization’s API keys for an AI provider, they may access private data and run up large usage bills. AI gateways can be used to protect API keys issued by AI providers.
Obstacles
  • Market requirements are maturing and expanding, and many vendor offerings are still in the early stages of development. Anticipate that there will be frequent upgrades to gateway solutions to keep up with evolving standards and protocols.
  • Solutions like API gateways, integration platforms and AI development platforms are offering complementary capabilities to AI gateways. As the market matures, expect to see consolidation into a single offering that supports multiple use cases.
  • Caching of semantic responses poses distinct challenges. The variability in natural language prompts and the nuances in LLM responses make it difficult to standardize caching strategies. As such, organizations may struggle to effectively implement caching measures.
  • Latency is critical in AI deployments. Because AI gateways can introduce latency, it is essential to evaluate the latency and performance implications of introducing the gateway into the architecture.
User Recommendations
  • Use an AI gateway implementing multiple LLMs or deploying AI agents that need to integrate with existing applications using AI-specific protocols. AI gateways provide centralized control, security, and cost management that is focused on AI-specific requirements that are not supported by standard API gateways.
  • Start with a robust proof of concept (POC) when adopting AI gateways and carefully evaluate products against technical and operational needs before full deployment. As AI requirements evolve rapidly, select vendors with a proven ability to adapt to new standards and regulations by prioritizing those with a clear innovation roadmap to avoid costly retrofits and ensure long-term flexibility.
  • For simpler scenarios, such as rate-limiting of AI APIs, evaluate your existing API gateway’s capabilities to act as an AI gateway, to see if they already are or can be upgraded, to support your requirements.
Sample Vendors
Cloudflare; Kong; Kosmoy; Lunar.dev; Portkey; Radiant; Traefik Labs; TrueFoundry
Gartner Recommended Reading

At the Peak

Developer Productivity Insight Platforms

Analysis By: Alec Pallin, Frank O'Connor
Benefit Rating: High
Market Penetration: 5% to 20% of target audience
Maturity: Emerging
Definition:
Developer productivity insight platforms provide software engineering leaders, developers and C-suite stakeholders visibility into how engineering teams use time and resources, their operational effectiveness, and the impact of AI on developer productivity through both qualitative and quantitative data. Software engineering leaders should use these platforms to drive improved productivity and value delivery.
Why This Is Important
Developer productivity insight platforms provide a comprehensive view of key engineering metrics, enabling software engineering leaders to assess and demonstrate team value through data-driven insights. These platforms help leaders identify, plan and track improvement opportunities by setting data-based baselines and realistic goals. Finally, they also support measurement of the ROI from AI initiatives by capturing their impact on developer productivity and business value.
Business Impact
CIOs and senior engineering leaders must demonstrate team value, improve productivity and understand the impact of AI on engineering outcomes. Organizations across industries require data‑driven insights to improve visibility, efficiency and alignment with business goals. These platforms act as a single source of truth for engineering process data, providing a unified, comprehensive and transparent view of engineering processes to demonstrate value.
Drivers
  • Growing need for evidence-based measures of value: Software engineering leaders face increasing pressure to demonstrate the value delivered by engineering teams and require objective, data‑driven measures to quantify and communicate that value.
  • Need for higher productivity: Gartner has observed sustained client interest in developer productivity over the past year, particularly in strategies to “do more with less.” Developer productivity insight platforms equip teams to drive their own productivity improvements by providing visibility into bottlenecks and productivity barriers.
  • Measuring AI’s impact on developer productivity: As organizations move from AI pilots to scaled adoption across the software development life cycle (SDLC) and AI agents, software engineering leaders must demonstrate measurable productivity gains and business returns. Developer productivity insight platforms support this effort by tracking AI’s impact over time through consistent metrics and visibility.
  • Benchmarking: Many organizations want to understand how their engineering performance compares with peers. Developer productivity insight platforms support cross‑industry benchmarking and comparative analysis based on aggregated client data.
Obstacles
  • Competition from DevOps platforms: Leading DevOps platforms are closing the capability gap by adding features that deliver a more complete view of the SDLC, with a focus on productivity and value.
  • Data confidentiality and security: Developer productivity insight platforms require access to sensitive and often proprietary data. Organizations may not wish to grant such access, particularly when platforms operate on public cloud infrastructure.
  • Integration effort: Collecting and normalizing data from homegrown or heavily customized engineering tools can prove challenging.
  • Perception of micromanagement: Teams may perceive data collection as monitoring, which can erode trust, reduce productivity and increase the risk of talent attrition.
  • Organization maturity: The cost and effort to adopt a developer productivity insight platform may not justify the value in organizations with immature engineering processes.
User Recommendations
  • Improve development team productivity and engagement by selecting developer productivity insight platforms that offer developer‑focused automation and qualitative feedback tools, as integrating analytics into developer workflows provides real‑time guidance for better outcomes.
  • Demonstrate the business value delivered by software engineering teams by choosing platforms that link strategic business outcomes to measurable, tactical engineering metrics.
  • Apply qualitative insights from developers to identify opportunities to improve the developer experience, focusing on areas of work and environment that developers consider important but rate low in satisfaction.
  • Use developer productivity insight platforms to track AI‑driven productivity gains across the software engineering organization.
Sample Vendors
Allstacks; Atlassian (DX); BlueOptima; Faros; Jellyfish; LinearB; Oobeya; Opsera; Plandek; Sleuth; Swarmia; Uplevel; Waydev
Gartner Recommended Reading

Model Context Protocol

Analysis By: Andrew Humphreys, David Pidsley
Benefit Rating: High
Market Penetration: 1% to 5% of target audience
Maturity: Emerging
Definition:
Model Context Protocol (MCP) is an open standard that enables two-way communication between AI models and other applications and data sources. It provides a standardized way for applications to share contextual information and expose tools and capabilities to AI systems that use LLMs.
Why This Is Important
MCP standardizes access to external data and tools, simplifying AI integration, improving interoperability, and reducing the need for custom code or non‑AI‑friendly APIs. published by Anthropic in November 2024 and donated to the Agentic AI Foundation in December 2025. It continues to evolve with new features and use cases. MCP has seen rapid adoption driven by widespread community adoption backed by AI providers including OpenAI, Google, and Microsoft adopting the standard.
Business Impact
MCP can lead to more context-aware AI systems and can reduce integration effort by standardizing how to give AI models access to contextually relevant and up-to-date data as well as enabling AI systems to take actions. However, poor implementation and ungoverned adoption have led to additional security concerns, as well as strategic questions about its usage. It also may be eclipsed by future standards that have greater openness and security.
Drivers
  • Unlike traditional APIs, MCP enables AI agents and applications to discover available tools, data sources, and capabilities at runtime rather than just in the original design and implementation. This makes it easier to give access to new data sources or allow new actions without retraining the model or significantly altering the core system, increasing the flexibility and agility of the AI solutions.
  • AI service providers and multiple application, analytics, decision intelligence platforms and middleware vendors have rapidly built and are offering prebuilt remote MCP servers in their applications and MCP support in tools. This further helps grow the MCP community adoption and ecosystem and simplifies adoption. This can lead to increased value and innovation as teams can reuse predefined MCP servers for new use cases without needing to build new MCP servers or custom integration points.
  • A key claimed driver is that productivity and effectiveness for AI agent development are increased because connecting to data sources and tools is simplified as it removes the need for custom, point‑to‑point integrations to be built for each agent and strengthens governance, security, and auditability. However, if you already have a strong approach to using APIs or other approaches to standardize how data and tools are accessed this claim may be overstated.
Obstacles
  • Poor design and implementation of MCP, especially for legacy systems or diverse technology stacks, can lead to suboptimal architectures and governance. Integrating MCP into existing infrastructure requires adherence to best practices and design principles.
  • Security is a common risk from poor implementation, particularly related to authorization to data and tools and to risks with data privacy and security. Ensuring that classified information is protected and access is appropriately controlled is crucial.
  • MCP is relatively new and has been evolving rapidly with many version updates in the last 15 months. Early adopters may face challenges with keeping implementations up to date with the latest version of the standard.
  • Most MCP clients disproportionately emphasize tool listing and tool calling, often ignoring or minimally supporting the rest of the protocol. This narrows MCP’s use to “function calling at scale,” undercutting its intended role as a general context interchange protocol.
User Recommendations
  • Establish a formal review process to assess MCP use‑case risk. Start with read‑only MCP tools for contextual data. Introduce tools that modify data or trigger actions only with strong identity, security, and governance controls.
  • Implement an MCP gateway to ensure MCP access is governed, observable, and auditable, supporting trust, compliance, and centralized policy enforcement.
  • Allocate budget to manage technical debt, giving teams time to track MCP’s rapid evolution, update implementations as standards change, and retain the option to exit if MCP loses its strategic value.
  • Strengthen security using centralized SSO (e.g., Okta, Active Directory, AWS IAM, Auth0). Use per‑user authentication instead of shared API keys, enforce least‑privilege RBAC for AI agents, require continuous re‑authentication for long‑running sessions, and apply human‑in‑the‑loop approvals for high‑risk actions such as sensitive data access or critical transactions.
Sample Vendors
Anthropic; Axiom; Cloudflare; Composio; Google; Microsoft; OpenTools; Stripe; Zapier
Gartner Recommended Reading

AI Application Development Platforms

Analysis By: Cary Pillers
Benefit Rating: High
Market Penetration: 5% to 20% of target audience
Maturity: Adolescent
Definition:
AI application development platforms provide tools, runtime services and workflows to design, build, evaluate, deploy and manage AI-embedded applications. These platforms include access to foundation models and the ability to ground and guardrail them. Additionally, these platforms provide full life cycle support for responsible AI development and use cases, such as building AI agents, assistants and multimodal AI applications.
Why This Is Important
AI engineers are being tasked with building new types of applications, such as AI agents and multimodal AI applications, as well as embedding AI into existing enterprise systems. AI application development platforms aid developers by providing a single platform across the life cycle of creating and embedding AI into applications. These tools enable enterprises to build more quickly than using a combination of open-source technology, prebuilt models and APIs across disparate providers.
Business Impact
AI application development requires precision, performance and responsible AI as a foundation to enterprise use. Building multiagent systems, integrating multimodal AI, enabling trust in AI responses and controlling costs will be focal points for 2026 and beyond. These platforms enable developers to use the right model, enable trust through guardrails and evaluations, and manage performance and costs at scale.
Drivers
  • AI-driven product development roadmaps: Although AI is driving these roadmaps, consistency is lacking across the developer skill sets, evaluation immaturity, orchestration gaps and open-source models and tools. This hampers the inclusion of AI into enterprise solutions at scale. Furthermore, AI may create unacceptable risk by introducing bias and hallucinations into enterprise workflows. Developers require tooling across the design, build, test and deploy workflow with a foundation of responsible AI features, which AI application development platforms provide.
  • AI agents: AI agents are autonomous or semiautonomous software entities that use AI techniques to perceive, make decisions, take actions and achieve goals in their digital or physical environments. AI agents will increase productivity across employees and customers, and offer transformation of business models. Multiagent systems require AI agents to coordinate their efforts to achieve their goals.
  • Multimodal AI apps: Multimodal AI applications use different input modalities (images, speech, text) to interpret and reason. This enables new experiences, such as document intelligence, clinical decision support, supply chain anomaly detection and knowledge work automation. AI avatar applications are created using various AI techniques like natural language processing, synthetic voice, computer vision and video translation, among others, to offer more human-like user experience.
  • AI assistants: AI assistants boost employee efficiency, minimize cognitive load, amplify problem solving, accelerate learning pace, foster creativity and maintain their state of flow. They leverage an orchestration framework to enable process-driven decision support, personalization, contextualization and domain-specific knowledge.
Obstacles
  • Specialized skills: Building and deploying AI agents, assistants and avatars require developers to use new technologies.
  • Monitoring and governing AI solutions: Coordination between AI agents, AI assistants and multimodal AI applications requires careful monitoring, governance and evaluations to ensure that the system behavior achieves intended goals.
  • Trust in AI: Despite improvements, unpredictable behavior, bias and weak evaluations persist, leading to limited trust in AI applications’ responses and decisions.
  • Grounding generative AI models: The need to provide context to generative AI models is challenging, requiring well-crafted RAG solutions.
  • Input and output rigor: Use guardrails to mitigate risks and obstacles like undesired input behaviors (some of which may be malicious) and outputs, such as hallucinations.
  • Pricing models: Usage-based pricing models for these platforms and the underlying infrastructure present a risk for businesses as the costs can accrue rapidly.
User Recommendations
  • Choose AI application development platforms that provide a unified ecosystem instead of assembling solutions from disparate vendors, language models and AI services.
  • Increase the likelihood of success in your AI strategy by enabling guardrails, running production pilots and applying measurable KPIs to assess the effectiveness of AI techniques.
  • Embed evaluations to ensure the model meets the required specifications during development and in production.
  • Ensure generative AI models are loosely coupled to adapt to the rapidly evolving technology.
  • Utilize AI application development platforms across common use cases, such as RAG for chatbots, AI avatars and AI agents, to embed AI throughout employee- and customer-facing technologies.
  • Leverage the responsible AI features within AI application development platforms, including guardrails, to protect your investment, brand and reputation while leveraging novel AI technologies.
Sample Vendors
Alibaba Cloud; Amazon Web Services; Google; IBM; LangChain; Microsoft; OpenAI; Palantir; Tencent; Volcano Engine
Gartner Recommended Reading

AI Engineering

Analysis By: Soyeb Barot, Haritha Khandabattu, Gary Olliffe, Chirag Dekate, Arun Chandrasekaran
Benefit Rating: Transformational
Market Penetration: 5% to 20% of target audience
Maturity: Early mainstream
Definition:
AI engineering is the discipline of designing, developing, delivering, operating, and governing tools and systems that use/deploy/apply AI to deliver business value. The discipline unifies DataOps, ModelOps, LLMOps, AgentOps and DevSecOps to create a coherent development, deployment, and operationalization framework for AI-based solutions.
Why This Is Important
AI engineering matters because most enterprises no longer struggle with creating isolated proofs of concept; they struggle with repeatable production delivery. The value with AI comes from turning fragile AI experiments into governed, reusable capabilities. Few organizations have built the data management, AI model management, agentic orchestration and DevSecOps foundations to build, maintain or operate portfolios of AI solutions — ceding competitive ground to those that can. Cross-collaboration across teams with diverse skills is needed to build composite AI solutions. Enterprises must establish consistent pipelines supporting the full scope of AI models and agents.
Business Impact
AI engineering provides speed and control — faster delivery of AI solutions, with the governance and stakeholder alignment needed to manage risk and cost at scale. It establishes process flows to ensure necessary cross-domain collaboration across IT and business with the development and maintenance of AI-based solutions. With defined AI engineering processes, it is possible to deploy AI solutions into production in a structured, repeatable model. Significant engineering, process and cultural challenges must be addressed as part of building and deploying composite AI solutions at scale, and cross-functional alignment between AI engineering and business is the defining challenge enterprises are paying to solve in 2026.
Drivers
  • Scaling from pilot to production requires operational discipline across the full AI life cycle — from data ingestion and model engineering to agent deployment — in a governed, repeatable architecture combining classical ML, LLMs, RAG, APIs and agents.
  • DataOps, ModelOps (including LLMOps), AgentOps and DevSecOps provide best practices for moving artifacts through the AI development life cycle. Standardization across data and model pipelines is accelerating the delivery of AI solutions, whether the approach is retrieval-augmented generation (RAG) or fine-tuning techniques alongside implementations with models built using diverse AI techniques.
  • AI engineering enables discoverable, composable and reusable data, AI artifacts (such as data catalogs, knowledge graphs, code repositories, reference architectures, feature stores and model stores), and agents across the enterprise technical architecture. These are essential for scaling delivery of AI solutions enterprisewide.
  • AI engineering is being driven by demand for agentic AI solutions. AI engineering teams are adapting existing software development life cycle practices into new agent development.
  • The shift to agentic AI solutions requires AI engineering to address multiagent orchestration, tool use, autonomous decision loops, and real-time inference — demanding new practices beyond those developed for single-model deployments.
  • Regulatory pressure, sovereignty (including the EU AI Act) and competitive urgency are forcing enterprises to invest in AI engineering as the operational backbone that makes AI auditable, governable, and continuously deliverable.
Obstacles
  • The hardest part with building AI solutions is usually not model building; it is operating-model and platform execution to scale from pilots to production. This is where failure occurs in enterprises’ AI-native building ambitions.
  • AI engineering requires simultaneous maturity across multiple domains (data, model, agent, platform and governance) which most enterprises cannot develop in parallel at the required pace.
  • It requires integrating full-featured solutions with an ecosystem of tools, enabling operationalization capabilities, to address enterprise architecture gaps with minimal functional overlap. These include gaps around extraction, transformation and loading data stores, feature stores, model repositories, and agent repositories, and ensuring observability, orchestration, and governance across the life cycle.
  • AI engineering requires cloud maturity and possible rearchitecting, or the ability to integrate data and AI model and agent pipelines across various deployment contexts.
  • Tooling fragmentation is a critical obstacle. Enterprises accumulate separate point tools for pipelines, model registries, observability, evaluation, and governance — creating integration debt that slows delivery and undermines auditability.
  • Talent scarcity is another obstacle which can compound the problem. Engineers who can operate across data engineering, DevOps, model development, agentic orchestration, and platform infrastructure simultaneously are rare and in high demand.
User Recommendations
  • Maximize business value from ongoing AI initiatives by establishing an AI engineering practice that streamlines data, models and implementation pipelines. Simplify data and analytics pipelines by identifying the capabilities required to operationalize end-to-end AI development platforms and build AI-specific toolchains.
  • Implement CI/CD practices for all AI artifacts — data pipelines, models, and agents — to enable continuous delivery, rollback, and auditability across the AI development life cycle.
  • Apply platform engineering principles to the establishment and sustainment of optimized tools and technology capabilities that support AI engineering across multiple solutions and initiatives. Avoid piecemeal technology selection. Think ecosystem and platform to enable teams across domains to build and deploy AI-based systems.
  • Leverage cloud service provider environments as foundational to build AI engineering. At the same time, rationalize your data, analytics, and AI portfolios as you migrate to the cloud.
  • Adopt a platform approach with GenAI by investing in centralized AI engineering tools for automation, governance, and use-case enablement across a broad set of AI models and cloud service providers.
  • Upskill data engineering and platform engineering teams to adopt tools and processes that drive continuous integration/continuous development for AI artifacts (e.g., data, models, agents).
Sample Vendors
Akka; Anyscale; Amazon Web Services; CoreWeave (Weights & Biases); DataRobot; Google; Microsoft; NVIDIA (OctoAI); OneReach.ai; TrueFoundry
Gartner Recommended Reading

AI Governance

Analysis By: Svetlana Sicular
Benefit Rating: High
Market Penetration: 20% to 50% of target audience
Maturity: Adolescent
Definition:
AI governance is the process of creating policies, assigning decision rights, and ensuring organizational accountability for risks and decisions across the AI life cycle. Enterprises make decisions on the appropriate, safe use of AI to achieve business outcomes within governance guardrails. AI governance addresses the predictive, generative, and increasingly autonomous (agentic) nature of AI to ensure responsible use and regulatory compliance.
Why This Is Important
Effective AI governance must support AI progress by balancing the business value of AI with proper oversight, where oversight is a strategic enabler, not a bureaucratic bottleneck. AI governance efforts must span three directions:
  • An operating model
  • Policies and controls
  • Enabling oversight technologies
Scaling AI from experimentation to production-grade systems without governance is ineffective and dangerous, particularly as AI shifts toward high-agency ecosystems.
Business Impact
AI governance, as part of the enterprise governance structure, establishes, monitors, and enforces AI guardrails to support AI progress and business value. It provides a common operating model and technical oversight for:
  • Applications, models, and agentic AI
  • Risk management, privacy, sovereignty, and regulatory compliance
  • Trust and transparency to support and secure AI adoption
  • The right data, technologies, and roles for the AI portfolio
  • Ethics, fairness, and safety to protect the business and its reputation
Drivers
  • According to the Gartner AI Maturity and Organizational Mandates for 2026 Survey, establishing governance frameworks and ethical guidelines for AI use is the top-ranked action organizations are taking to reduce risks associated with AI systems. This keeps AI governance in the Peak of Inflated Expectations.
  • The rise of agentic AI: The transition from passive tools to autonomous, multiagent ecosystems introduces complex compounded risks, necessitating explicit accountability, guardrails, and continuous monitoring to prevent loss of control and unpredictable emergent behaviors.
  • Board mandates and regulatory pressure: CEOs and boards demand structured governance frameworks to assure AI value, safety, and sovereignty. Compliance with new laws (e.g., EU AI Act), emerging AI insurance offerings, and the need for rigorous audit trails are unlocking dedicated governance budgets.
  • Data sensitivity and privacy: Utilizing proprietary, sensitive data in AI models strains organizational trust. The aggregation problem, where isolated safe data becomes sensitive when merged, requires governance frameworks beyond traditional data management.
  • The need for strategic enablement: Organizations are moving away from ad hoc pilots toward centralized or hybrid (federated) operating models to ensure interoperability, scale AI safely, and replace shadow AI with preapproved, secure pathways.
  • Workforce trust and AI literacy: Achieving widespread AI adoption requires addressing workforce skepticism by establishing transparent governance, promoting critical thinking about AI’s probabilistic nature, and defining clear verification processes.
  • Continuous monitoring for risk and value: The dynamic nature of AI demands monitoring and observability tools to track drift, bias, and compliance deviations across the entire AI life cycle.
Obstacles
  • The conflict between innovation and safety often causes employees to view governance as a bottleneck, leading to disconnected practices.
  • Lack of clarity in decision rights for tools, models, software, and principles leads to suboptimal operating models.
  • Outdated governance practices impair AI guardrails and scaling AI.
  • Many organizations do not balance AI value assurance and AI risk management, ignoring the former.
  • Interoperability, reliability, and evaluation in agentic AI environments, if ignored, could impede complex workflows. Compounded agentic AI risks due to complexity in orchestration are more likely to cause failures than an individual component.
  • The nondeterministic nature of AI and the opaque reasoning make continuous evaluation, reliability testing, and establishing a chain of liability difficult.
  • Technologies supporting AI governance remain fragmented, often focusing on isolated capabilities like evaluation or security, rather than offering unified, systemic oversight.
User Recommendations
  • Focus governance on your AI portfolio; don’t “boil the ocean.” Pace AI governance with AI speed.
  • Adopt a centralized governance operating model initially for stability, and transition to a federated model to balance central control with decentralized agility as AI maturity grows.
  • Establish and refine processes for making AI-related decisions.
  • Define levels of use-case criticality to focus AI governance on what matters the most and allow freedom for innovation.
  • Explicitly define accountability, decision rights, and an AI agent code of conduct to ensure autonomous agents and human stakeholders operate within clear ethical and regulatory guardrails.
  • Invest in AI literacy to proactively increase the quality of AI-related decisions.
  • Mandate human-in-the-loop oversight and define escalation procedures to intervene in critical, high-stakes decisions and manage the compounded risks of multiagent orchestration.
  • Implement tools for AI oversight, review, validation, and evaluation.
Gartner Recommended Reading

Curated OSS Catalogs

Analysis By: Aaron Lord
Benefit Rating: High
Market Penetration: 5% to 20% of target audience
Maturity: Emerging
Definition:
Curated open-source software (OSS) catalogs are trusted repositories of vetted OSS dependencies that enable policy-based curation. The repositories can either be curated externally by a provider or internally by the organization based on security, compliance and operational policy checks. These checks should include open severe vulnerabilities, license compliance and project health to prevent security, legal and viability risks.
Why This Is Important
OSS usage is common for application development. Yet, risks to the software supply chain from malicious and vulnerable OSS dependencies continue unabated. Providing developers with a curated catalog of trusted open-source artifacts improves governance without impeding developer experience. It ensures centralized visibility, streamlines patching and enables versioning.
Business Impact
Curated catalogs can help every software engineering organization reduce exposure to OSS security risks. By centralizing the responsibility for security scanning, dependency management and compliance monitoring, curated catalogs allow organizations to benefit from open-source innovation while mitigating risks. This makes it easier to ensure regulatory compliance, particularly when creating software bills of materials (SBOMs) or monitoring common vulnerabilities and exposures (CVE) reports.
Drivers
  • The onslaught of software supply chain attacks on upstream OSS is making security and risk teams concerned about unfettered use of OSS within the organization. These concerns include security risks such as typosquatting, backdoors, dependency confusion and zero-day vulnerabilities, as well as operational risks from using out-of-date and unmaintained packages, which are easy targets for software supply chain attacks.
  • Government regulations are mandating the use of SBOMs, necessitating greater visibility and auditability of OSS being used internally within organizations. For organizations using open-source components in their technology stacks, these curated catalogs offer a path toward standardized, secure, compliant and maintainable software ecosystems.
  • Platform engineering teams provide curated development experiences via secure “paved roads” for a broad range of software development life cycle (SDLC) use cases. Therefore, providing a curated OSS catalog aligns well with platform engineering goals and helps meet the need for both developer autonomy and governance.
  • A few vendors now provide trusted OSS packages “as a service,” with the assurance that the curated catalog contains prevetted OSS packages that are safer to use. This delivery model helps drive adoption by making it easier to create a catalog. The vendors assume responsibility for finding and fixing vulnerabilities, as well as for blocking packages that pose software supply chain security risks.
  • Usage of AI coding agents is creating more code than ever before. However, they can be directly integrated with curated OSS catalogs. This will ensure that agents only utilize OSS artifacts that have been curated into the application development environment.
Obstacles
  • Curated OSS catalogs deliver maximum benefit when they are part of a DevOps platform consumed by multiple product teams. Therefore, organizations that lack a platform engineering approach may find it difficult to scale the use of curated catalogs.
  • Developers may perceive a curated catalog as a threat to their freedom and autonomy, especially if it is positioned as an “enforcer” of OSS governance rather than an “enabler.”
  • The lack of integration with continuous integration/continuous delivery (CI/CD) pipelines can make it difficult to systematically block packages from deploying to production using automated verification. This shortcoming dilutes the purpose of the catalog and undermines its utility.
  • Receiving investments for a curated OSS catalog may be difficult if the organization lacks awareness of the risks of using OSS. Procurement of buy-in and funding for curated OSS catalogs may require significant upfront research to justify its costs.
User Recommendations
  • Establish and maintain a “trusted repository” of vetted and approved dependencies aligned with your governance policy, including supporting processes.
  • Prevent vulnerable, unapproved, malicious and noncompliant packages from entering the SDLC by continuously scanning and assessing packages in the curated catalog for security, compliance and operational risks.
  • Treat the curated OSS catalog as a dynamic repository rather than a static list. Establish trust in the OSS based on continuous policy checks, rather than a single point-in-time scan.
  • Ensure that the curated catalog integrates with developer workflows and existing CI/CD pipelines to automate policy enforcement checks.
  • Prefer OSS catalogs that offer automated recommendations for alternative packages or versions of blocked software to improve developer experience. AI assistants typically drive this feature to provide a conversational interface for developers.
Sample Vendors
ActiveState; Chainguard; Docker; Echo; JFrog; Lineaje; Red Hat; Seal Security; Sonar; Sonatype
Gartner Recommended Reading

FinOps

Analysis By: Lydia Leong
Benefit Rating: High
Market Penetration: 20% to 50% of target audience
Maturity: Early mainstream
Definition:
FinOps represents the integration of agile, lean and DevOps practices into a form of IT financial management. Implementations are optimized to manage, in near real time, the dynamic life cycle of metered-by-use services such as cloud. It is a cultural practice emphasizing cross-functional collaboration. Its many stakeholders reside within the business as well as IT, and there is distributed accountability for its success.
Why This Is Important
Digital democratization frequently empowers the business to buy its own IT solutions or to consume services on-demand. This requires adaptive cost governance, agile budgeting, and management of cost complexity and variability. FinOps helps organizations optimize metered costs, extract greater value from investments and connect cost to business value. FinOps emerged out of the need for cloud financial management (CFM), but organizations may also aspire to apply it to traditional IT solutions.
Business Impact
FinOps emphasizes a shift from ad hoc cost hygiene activities to a cross-functional continuous life cycle for metered consumption. It is focused on maximizing the value to the business rather than minimizing expenses, including making conscious trade-offs. For example, business leaders may reasonably make the informed decision to spend more to deliver a better user experience or to ignore cost-related technical debt so that application teams can focus on delivering more features.
Drivers
  • Public cloud costs are of concern to many customers. Some customers experience unexpectedly high cloud costs because they have fulfilled a far greater business demand for cloud services than the IT organization had forecast.
  • Ungoverned cloud adoption leads to uncontrolled and careless spending. Organizations without an effective CFM practice cannot properly plan, track or optimize their cloud costs.
  • Services with metered consumption have costs that can change continuously, and therefore demand near-real-time cost management that can handle the extreme granularity and complexity.
  • Digital democratization has shifted the decision makers for consumption, requiring a new accountability model to ensure that cost is factored into decisions. Cost-efficient solutions require designing cost-aware architectures.
  • Public cloud services enable unprecedented cost transparency and thus better correlation between IT costs and business value. However, organizations that do not effectively manage cloud economics cannot assist the business in making thoughtful decisions about the top line versus the bottom line in cloud-enabled digital products and services.
  • The FinOps Foundation (FOF), founded in 2019 by several cloud cost management tool vendors, created and popularized the use of the FinOps term. Its membership has grown over time and now includes many vendors that sell FinOps tools and services (which FOF will badge as FinOps Certified Platforms or Services), numerous cloud providers, several global system integrators and some enterprises (mostly financial services institutions). These members promote FinOps concepts to their users, increasing the hype.
  • Organizations that have been successful with CFM increasingly want to apply FinOps practices to internal IT services. Vendors are promoting the concept of converging FinOps, IT financial management and software asset management tools. However, the market landscape remains immature and the tools are not well integrated.
Obstacles
  • Not all organizations have a business case for a full CFM practice, although almost all organizations that have meaningful cloud adoption need to perform cost management. Tool and labor costs needed for a FinOps program may exceed its cost savings.
  • CFM needs to be a cross-functional and cultural practice, not solely the responsibility of a dedicated FinOps team (which many organizations cannot cost justify).
  • Application teams and the business owners of applications need motivation to optimize their cloud spend. This usually requires coupling showback to incentives (or penalties) or performing chargeback, forcing these entities to be accountable for what they spend.
  • Although FinOps tools can often recommend optimizations for cloud infrastructure, they have limited capabilities for platform and other higher-level services — especially AI services.
  • The experimental nature of many AI initiatives challenges FinOps teams, who often expect predictable budgets and timelines. However, premature optimization is a waste of time.
User Recommendations
  • Establish a cross-functional CFM practice — not just a FinOps team — once public cloud IaaS and PaaS adoption has proceeded beyond the pilot stage.
  • Assign ownership of the CFM practice to your cloud center of excellence (CCOE) or other cloud governance function, which will collaborate with sourcing, procurement and vendor management, and IT finance as needed. The CCOE should set the policies and guidance, but cloud operations teams are typically responsible for implementing CFM tools and assisting application teams with optimizations.
  • Design cost-aware architecture when custom-developing cloud solutions. Application design and implementation will be the primary influence on cloud costs.
  • Use performance-based rightsizing and accept the smallest recommended size by default when migrating existing applications to the cloud through a lift-and-shift (that is, a rehost) approach. Most organizations tend to oversize virtual machines when they migrate due to unwarranted concerns about performance and “headroom.”
Sample Vendors
Broadcom (VMware); CloudBolt; Flexera; IBM; Umbrella
Gartner Recommended Reading

AI Agents

Analysis By: Tom Coshow, Haritha Khandabattu
Benefit Rating: High
Market Penetration: 5% to 20% of target audience
Maturity: Early mainstream
Definition:
AI agents are autonomous or semiautonomous software entities that use AI techniques to perceive, make decisions, take actions and achieve goals in their digital or physical environments.
Why This Is Important
AI agents have the ability to make decisions and take action in their target environment to achieve organizational goals. By using AI practices and techniques such as LLMs, organizations are creating and deploying AI agents to achieve complex tasks.
Business Impact
AI agents have the potential to:
  • Revolutionize a broad range of industries and environments with their ability to automate tasks from consumer, industrial, data analytics, content creation and logistics.
  • Make informed decisions and interact intelligently with their surroundings.
Drivers
  • Generative AI breakthroughs: Reasoning models and LAMs advance the ability to plan a complex series of actions.
  • Multimodal understanding: The ability to use diverse modalities like vision, audio and language enables more general and flexible AI agents. This allows automatic adaptation to changes in the workflow, user interface or API. With this, one can create advanced workflows without explicit programming, significantly reducing the development time and effort for automation.
  • Increased decision-making complexity: AI is increasingly used in real-world engineering problems containing complex systems, where large networks of interacting parts exhibit emergent behavior that cannot be easily predicted. AI agents can learn, plan and execute in complex environments.
  • Composite AI, including neurosymbolic models: Advances in models that improve planning and problem solving are enabling more complex AI agents. AI agents can utilize a wide variety of AI practices to forecast, make decisions and plan.
Obstacles
  • Vulnerabilities: Due to the complexity of the AI agent system, all components face various potential vulnerabilities such as access security, data security and governance.
  • Lack of trust: Users are unsure whether they can trust the technology to accurately predict and execute tasks independently. Without a human in the loop, agents may take multiple consequential actions in rapid succession and bring about significant impacts before a human notices.
  • Interpretability and oversight: Action policies may be opaque and have poor explainability, requiring mechanisms for human interpretability, oversight and control.
  • Pace of change: The technologies deployed, from models to tooling, and the framework options available are changing rapidly, making it difficult for organizations to define their roadmap.
User Recommendations
  • Incorporate AI agents into strategic planning by investing in understanding their capabilities and potential applications in various environments, considering their increasing autonomy and wide-ranging usability.
  • Investigate the possibilities of utilizing multiagent systems, collectives of AI agents, that can operate both collaboratively and independently, enhancing adaptability and flexibility in response to different tasks and scenarios.
  • Promote the development and integration of the use of a variety of AI practices, enabling learning, negotiation and decision-making capabilities.
Sample Vendors
Amazon; Anthropic; CrewAI; Google; LangChain; Maisa; Microsoft; OneReach.ai; OpenAI; Salesforce
Gartner Recommended Reading

Cluster Fleet Management

Analysis By: Tony Iams
Benefit Rating: High
Market Penetration: More than 50% of target audience
Maturity: Adolescent
Definition:
Cluster fleet management addresses gaps in managing Kubernetes at scale and helps avoiding lock-in. Key capabilities include deployment and upgrading of orchestration software, and distribution of containerized applications and operational policies across clusters. Offerings can support single or heterogeneous Kubernetes distributions.
Why This Is Important
As Kubernetes adoption grows within an organization, infrastructure and operations (I&O) teams increasingly find they must deploy and manage multiple clusters that may be located on-premises, in the cloud or at the edge. An organization’s approach to managing fleets of clusters can impact hybrid and multicloud initiatives, the resilience of containerized workloads and the design of developer workflows for delivering containers.
Business Impact
Scaling to meet customer demand often requires leading-edge, cloud-native technologies. Because many cloud-native technologies utilize containers and Kubernetes, cluster fleet management is an important technical consideration for organizations seeking to sustain or accelerate growth in digital products.
Drivers
  • Modernization: Most enterprises have adopted containers within one or more projects to enable the building and/or refactoring of applications. Many I&O teams find that, as deployments and initiatives increase, they must support a broader set of operational requirements.
  • Hybrid and multicloud container initiatives: Most enterprises realize that separate Kubernetes clusters will need to be deployed — for a variety of technical and business requirements. Depending on the deployment, it may not be possible — or preferential — to use the same Kubernetes distribution in all environments. With a multidistribution (heterogeneous) approach comes the potential for reduced lock-in, but also added complexity and the need to ensure consistency in management across multiple distributions and/or cloud container services.
  • Multitenancy: Where separate clusters are deployed in one location for different teams or classes of applications, by deploying separate clusters for teams or applications, the “blast radius” can be limited if a catastrophic outage makes an entire cluster unresponsive. In the cloud, assigning separate clusters to different teams also simplifies cost showback and chargeback.
  • Disaster recovery and/or resilience: Where clusters are maintained in separate data centers to take over when a primary data center becomes inaccessible.
  • Edge deployments: Where containers need to run in many remote locations, each of which requires a separate lightweight Kubernetes cluster deployed on minimal hardware.
Obstacles
  • Lack of standards: Kubernetes does not currently define a standard approach for distributing its functionality across multiple clusters, and open-source projects for extending clusters are not yet mature. Managing cluster fleets requires proprietary commercial solutions, which introduces concerns about operational lock-in and vendor viability (at least until a standard approach is defined).
  • Cost: Some public cloud services charge separately for each Kubernetes control plane, which increases costs if multiple clusters are deployed.
  • Complexity: Deploying containerized applications to fleets of clusters based on heterogeneous Kubernetes distributions adds complexity. Managing the state of multiple clusters requires mature Kubernetes operational practices such as GitOps and platform engineering. Heterogeneous fleet management tools are limited in the range of operations they support. Developer workflows may not be sufficiently flexible to account for operational differences in each distribution.
User Recommendations
  • Determine whether to deploy homogeneous cluster fleets based on a single distribution, or to take a platform-engineering approach to normalize operations and application deployment workflows across clusters based on different Kubernetes distributions.
  • Select a management platform to control the life cycles and operation of multiple Kubernetes clusters and to automate the distribution of Kubernetes resource definitions, policies and identities across multiple clusters.
  • Automate cluster life cycle management as much as possible, using an immutable approach to define abstract classes of clusters that can be used to launch specific cluster instances on demand.
  • Assess advanced cluster deployment approaches such as virtual clusters and hosted control planes to optimize the cost of infrastructure when deploying multiple clusters.
  • For organizations with a high level of DevOps maturity, use distributed GitOps for deploying applications, policies and other extensions across multiple clusters.
Sample Vendors
Avesha; Broadcom; Google; Microsoft; Mirantis; Rafay Systems; Rancher; Red Hat; Spectro Cloud; SUSE
Gartner Recommended Reading

Sliding into the Trough

Developer Experience

Analysis By: Alec Pallin, Brian Minning
Benefit Rating: High
Market Penetration: 20% to 50% of target audience
Maturity: Early mainstream
Definition:
Developer experience refers to all aspects of interactions between developers and the tools, platforms, processes and people they work with to develop and deliver software products and services. A superior developer experience delivers an environment in which developers can do their best work with minimal friction and maximum flow.
Why This Is Important
Software development teams work in an increasingly complex environment with a growing array of tools, technologies, architectures and processes across the software delivery life cycle. This complexity results in friction and increased cognitive load for developers, limiting their ability to deliver value. Developer experience initiatives holistically address the causes of friction and frustration for teams, enabling them to focus on the highest-value activities with minimal distraction.
Business Impact
A superior developer experience drives a number of key organizational outcomes. Analysis of data from Gartner’s Developer Experience Assessment reveals developers with a high-quality developer experience are more likely to:
  • Achieve their target business outcomes, such as revenue growth and user satisfaction.
  • Have higher productivity, including better delivery flow, speed to market, release cadence and delivery predictability.
  • Have high intent to stay with their current employer.
Drivers
  • Pressure to boost developer productivity with AI: Software engineering leaders face intensifying pressure to increase their teams’ productivity with the emergence of AI tools. AI tools, such as AI coding assistants and agents, promise significant productivity gains for developers. However, fully realizing these gains requires leaders to align tool selection and use cases with developer experience needs and priorities. Partner with developers to identify friction points within their workflows where AI can have the greatest impact on their efficiency.
  • Growing complexity of software architectures and development technologies: AI-driven changes to how developers work, such as prompt engineering and agentic frameworks, are raising developers’ cognitive load. To deliver a superior developer experience and enable teams to do their best work, leaders must take steps to manage complexity for their teams, such as by providing them with internal developer portals, platforms and AI tooling.
  • Improve software quality: Software engineering leaders are under mounting pressure to elevate code quality and reduce defects. Organizations with stronger developer experience more consistently meet software quality and security standards, and require investments in tools and capabilities that help developers protect reliability while sustaining velocity.
Obstacles
  • Securing stakeholder buy-in for developer experience investments: Developer experience initiatives risk underinvestment unless stakeholders see a strong business case. Failure to secure support erodes developer trust and future participation in developer experience improvement efforts.
  • Unclear ownership of developer experience initiatives: With the end-to-end developer journey spanning a complex mix of processes, technologies and stakeholders, organizations without well-defined accountability for managing developer experience initiatives risk piecemeal and overly narrow approaches to addressing needs.
  • Incomplete understanding of challenges: The developer experience entails more than just tools; it also includes other aspects of the end-to-end developer journey such as onboarding, upskilling and workflow design. Organizations that fail to gain a holistic view of development teams’ leading pain points risk focusing narrowly on lower-impact improvements.
User Recommendations
  • Treat the developer experience like a product. Assign clear accountability for delivering the end-to-end developer journey to a dedicated developer experience team and product owner, while enabling individual development teams to identify and address local developer experience needs.
  • Regularly gather feedback from development teams about their leading challenges and pain points. Employ quantitative and qualitative methods, such as surveys, developer journey-mapping workshops and tool telemetry, and use this data to identify the highest-priority developer experience improvement opportunities.
  • Track developer experience improvements over time using business and team well-being metrics, and correlate changes with delivery outcomes to verify impact.
Sample Vendors
Allstacks; Culture Amp; DX; LinearB; Opsera; Swarmia
Gartner Recommended Reading

Agentic Cloud Development Environments

Analysis By: Manjunath Bhat
Benefit Rating: High
Market Penetration: 5% to 20% of target audience
Maturity: Adolescent
Definition:
Agentic cloud development environments (CDEs) offer a secure and governed approach to software development, providing either a cloud-hosted or self-hosted workspace equipped with agentic coding tools. These sandboxed, isolated environments separate the development workspace from physical endpoints, which simplifies administration and reduces the risk of running AI-generated untrusted code. In an agentic CDE, both human engineers and AI agents work together on code artifacts.
Why This Is Important
As AI coding agents advance from completing function snippets to developing complete functionalities, local development environments become a liability due to compute constraints, security risks (like privilege escalation), and the difficulty in maintaining consistency and reproducibility. Agentic CDEs resolve these issues by offering a secure, automated workspace that is optimized to meet the computational and experiential needs of humans and software engineering agents.
Business Impact
Regulated organizations such as banks and government agencies stand to gain the most. These organizations can meet regulatory and security obligations by protecting intellectual property and securing sensitive data. Organizations building AI applications can use agentic CDEs to benefit from hosted GPUs for faster training and inference. Software engineering leaders benefit from faster onboarding with new hires gaining same-day access to a consistently reproducible development environment.
Drivers
  • AI-native software engineering: CDEs are increasingly used as a distribution channel for developers to adopt approved agentic developer tooling, such as AI coding and augmented testing agents. This drives the adoption of agentic CDEs while simultaneously minimizing the security and compliance risks associated with using unvetted AI tools within the organization.
  • Mitigating software supply chain security risks: Security teams gain centralized control to manage, govern and secure development environments, significantly reducing the threat of supply chain attacks. This includes the ability to implement zero-trust access policies for source code and the delivery pipeline. Furthermore, because CDEs are defined by code, development tools and dependencies can be kept consistent.
  • Improving developer experience: The complexity and unreliability associated with maintaining and configuring local machines, due to numerous plug-ins, MCP and API integrations, is drastically reduced. Agentic CDEs provide seamless integration with Git repositories and continuous integration/continuous delivery (CI/CD) tools.
  • Reducing configuration drift: Developing and testing modern, complex applications such as mesh services, event-driven apps and container-native applications is challenging on local machines. Agentic CDEs make it easier for developers to ensure consistency of development environments.
  • Protecting intellectual property (IP) in outsourcing: Organizations increasingly use CDEs as a primary option to secure their IP when outsourcing software development. Gartner client inquiries indicate IP protection and risk mitigation as the primary demand drivers for CDEs among regulated companies.
Obstacles
  • Agentic CDEs introduce additional costs on top of existing DevOps tooling expenses. These increased costs can be excessive, particularly for development teams that currently leverage open-source tools for application development and delivery on local systems.
  • Agentic CDEs must evolve to keep pace with innovations in the agentic AI technology landscape. For example, IT administrators will require these tools to provide AI cost transparency and enforce necessary guardrails.
  • Security and compliance policies may prohibit using the public cloud for development, which could rule out agentic CDEs that depend on public cloud services. It is important to note, however, that CDEs can be provisioned in self-hosted environments such as a private data center.
  • Developers may resist CDEs because it could hinder their capacity for experimentation and rapid innovation due to restrictive IT governance policies.
User Recommendations
To effectively integrate agentic CDEs into development workflows, software engineering leaders should:
  • Prioritize sandboxing and isolation capabilities to prevent misbehaving or compromised agents from accessing sensitive data, restricted network segments, or local servers and endpoints. Enforce privilege access management controls and network policies to minimize the risk of data exfiltration.
  • Implement AI governance and security controls to reduce shadow AI risks and compliance gaps. Deploy an AI gateway to mediate traffic between the CDE and model providers, and consolidate API keys, create audit logs and provide visibility for token consumption. Integrate CDEs with enterprise secret managers (like HashiCorp Vault) to ensure agents can access necessary credentials.
  • Take a platform engineering approach to defining development environments as code — tools, libraries and configurations to support the applications being built and underlying technology stacks.
Sample Vendors
Anysphere (Cursor); Citrix; Coder; Daytona; Docker; E2B; Harness; Microsoft; Ona; Together AI
Gartner Recommended Reading

GitOps

Analysis By: Paul Delory, Arun Chandrasekaran
Benefit Rating: High
Market Penetration: 20% to 50% of target audience
Maturity: Adolescent
Definition:
GitOps is a closed-loop control system for cloud-native applications. According to the canonical OpenGitOps standard, the state of any system managed by GitOps must be expressed declaratively, be versioned and immutable, be pulled automatically, and be continuously reconciled. The term “GitOps” is often used more expansively, usually as shorthand for automated operations or continuous integration/continuous deployment (CI/CD), but this is incorrect.
Why This Is Important
GitOps can be transformational. GitOps workflows deploy a verified and traceable configuration (such as a container definition) into a runtime environment, bringing code to production with only a Git pull request. All changes flow through Git, so they are version-controlled, immutable and auditable. Developers interact only with Git, using abstract, declarative logic. GitOps extends a common control plane across Kubernetes (K8s) clusters, which is increasingly important as clusters proliferate.
Business Impact
By operationalizing infrastructure as code, GitOps enhances the management and resilience of services:
  • GitOps can improve version control, automation, consistency, collaboration and compliance.
  • Resource declarations are version-controlled, modular and stored in a central repository, making them easy to reuse, verify and audit.
  • Resource declarations are readily consumable by AI agents, which can automate PRs and drift detection.
Drivers
  • Kubernetes adoption and maturity: GitOps must be underpinned by an ecosystem of technologies, including tools for automation, infrastructure as code, CI/CD, observability and compliance. Kubernetes has emerged as a ready-made foundation for GitOps, because the continuous reconciliation loop at the heart of K8s complements the GitOps model. As Kubernetes adoption grows within the enterprise, GitOps can, too.
  • Need for increased deployment velocity and agility: Speed and agility of software delivery are critical metrics that CIOs care about. As a result, IT organizations are pursuing better collaboration between infrastructure and operations (I&O) and development teams to drive shorter development cycles, faster delivery and increased deployment frequency. GitOps is the latest way to drive this type of cross-team collaboration.
  • Need for increased reliability: Speed without reliability is useless. The key to increased software quality is effective governance, accountability, collaboration and automation. GitOps can enable this through transparent processes and common workflows across development and I&O teams. Automated change management helps to avoid costly human errors that can result in poor software quality and downtime.
  • Talent retention: Organizations adopting GitOps have an opportunity to upskill existing staff for more automation- and code-oriented I&O roles. This allows staff to learn new skills and technologies, increasing employee satisfaction and retention.
  • Cultural change: By breaking down organizational silos, development and operations leaders can build cross-functional knowledge and collaboration skills across their teams to enable them to work effectively across boundaries.
  • Cost reduction: Automating infrastructure eliminates manual tasks and rework, improving productivity, which can contribute to cost reduction.
  • Compliance requirements: The declarative nature of GitOps leaves an easy audit trail for software changes, improving compliance.
Obstacles
  • Prerequisites: GitOps is only for cloud-native applications. Many GitOps tools and techniques assume the system is built on Kubernetes. By definition, GitOps requires software agents to act as listeners for changes and help implement them. GitOps is possible outside of Kubernetes; however, in practice, K8s will almost certainly be used. Thus, GitOps is necessarily limited in scope.
  • Cultural change: GitOps requires a cultural change that organizations must invest in. IT leaders must embrace process change. This requires discipline and commitment from all participants to do things differently.
  • Skills gaps: GitOps requires automation and software development skills, which many I&O teams lack. Practitioners and organizations must be mature enough to operate infrastructure through code. Many are not.
  • Organizational inertia: GitOps requires collaboration among different teams, which requires mutual trust to be successful.
User Recommendations
  • Target cloud-native workloads initially: Your first use case for GitOps should be operating a containerized, cloud-native application that is already using both Kubernetes and a continuous delivery platform, such as Flux or Argo CD.
  • Build an application delivery platform: This is the foundation of your GitOps efforts. Your platform should manage the underlying infrastructure and deployment pipelines, while enforcing security and policy compliance.
  • Embed security into GitOps workflows: Security teams must shift left, so the organization can build holistic CI/CD pipelines that deliver software and configure infrastructure, with security embedded in every layer.
  • Be wary of vendors trying to sell you GitOps: GitOps isn’t a product you buy. It is a workflow and a mindset shift that becomes part of your DevOps culture. Tools that expressly enable GitOps can be helpful, but GitOps can be done with nothing more than standard continuous delivery tools that support Git-based automation.
Sample Vendors
Akuity; GitLab; Harness; Red Hat; Upbound
Gartner Recommended Reading

Self-Service Environment Management

Analysis By: Chris Saunderson
Benefit Rating: High
Market Penetration: 5% to 20% of target audience
Maturity: Adolescent
Definition:
Self-service environment management tools empower engineers to provision and manage their own environments for building and delivering software. Platform teams use these tools to create environment blueprints and codify governance policies using infrastructure-as-code and policy-as-code capabilities. Product teams consume the environment blueprints and benefit from automation, built-in security, compliance and cost guardrails, and improved insight into application behavior.
Why This Is Important
In the absence of a self-service, automated approach to environment management, every environment becomes a uniquely managed entity, especially in cloud-native environments. This strategy can improve the developer experience by providing consistent, reproducible environments that minimize cost overruns, while increasing security and reducing compliance risks. The absence of ready-to-use environments makes it challenging to continually test software quality or to add new features (such as AI enablement).
Business Impact
Organizations can realize business value from self-service environment management tools in three key ways:
  • Improve developer productivity and experience by minimizing wait times on other teams, increasing overall business agility.
  • Ensure consistency and standardization by codifying governance policies and environment creation processes.
  • Lower the barrier for testing applications at runtime for functional and nonfunctional requirements, which improves reliability.
Drivers
  • The widespread adoption of platform engineering practices, such as “paved roads” and guardrails as a focus on improving developer productivity, has been a major contributing factor to the acceptance of self-service environment management tools.
  • Environment management is one of the key capabilities of internal developer platforms that help developers self-serve secure, compliant, and ready-to-use environments on demand. Environments can be managed through a life cycle approach to automatically shut down and not incur costs when not in use.
  • Containerized application delivery requires solid environment management to prevent friction, especially with the sprawl of Kubernetes clusters to host different environments across development, test, staging, and production.
  • The use of “as-code” approaches to infrastructure automation, policy automation, and configuration management enables the separation of concerns through self-service-based, templatized workflows. Platform teams create environment templates to be consumed by product teams.
  • Environments are a logical abstraction spanning applications, infrastructure, configuration, policies, and core services (such as DevOps pipelines, database, and service mesh). Providing environments “as-a-service” to internal developers makes it easier to abstract away undifferentiated complexity and enable consistent governance, as well as minimize cognitive load.
  • Cost optimization requires visibility into resource use. Self-service environment management ensures resources are consumed only when needed, minimizing waste and leading to significant cost savings, and can be enforced through environment life cycle capabilities.
Obstacles
  • Effective use of self-service environment management requires a consumer-focused approach that is created by platform teams curating a frictionless developer experience by providing “environment-as-a-service” to internal development teams. Without a platform team or a consumer focus, environment provisioning typically takes a backseat, goes through a service desk, and requires many manual steps involving multiple approvals.
  • The lack or scarcity of mature infrastructure automation skills in many organizations poses another bottleneck. Enabling developers with self-service access to environments requires enabling environments as code at a higher level of abstraction than just infrastructure components (e.g., GitOps and reconciliation patterns).
  • The use of legacy infrastructure and middleware components lacking API access makes them harder to codify, making developer self-service for environment management more difficult.
User Recommendations
  • Establish platform teams to improve the developer experience of environment provisioning using self-service environment management tools. This approach requires all infrastructure to be represented in code following infrastructure-as-code (IaC) principles and practices.
  • Enable guardrails using policy as code as part of environment provisioning templates to minimize risk exposure (e.g., credentials and secrets) and enforce cost-governance controls. Enable configuration drift detection with support for overrides to allow approved deviations from baseline environment templates.
  • Enable developers to re-create the state of any environment from information stored in version control and use the same deployment process for every environment, including production. This tactic helps ensure deployments are repeatable and, in the event of a disaster, you can restore the state of production in a deterministic way.
Sample Vendors
Appvia; Cycloid; Facets.cloud; Humanitec; Massdriver; Okteto; Qovery; Quali; Resourcely; Syntasso
Gartner Recommended Reading

Site Reliability Engineering

Analysis By: George Spafford, Daniel Betts
Benefit Rating: Transformational
Market Penetration: 5% to 20% of target audience
Maturity: Adolescent
Definition:
Site reliability engineering (SRE) is a collection of systems and software engineering principles used to design and operate scalable, resilient systems. Site reliability engineers collaborate with customers or product owners, using operational data and feedback, to define service-level indicators (SLIs) and service-level objectives (SLOs). Site reliability engineers work with product or platform teams to design, operate and continuously optimize systems that meet defined SLOs.
Why This Is Important
SRE is vital because it shifts reliability from an “IT problem” to a business strategy. By using reliability-focused metrics, it provides a data-driven framework to balance fast feature delivery with system stability. SRE drives efficiency by aggressively automating manual “toil,” allowing teams to scale without linear headcount growth. Ultimately, SRE ensures that reliability is a shared responsibility, protecting the customer experience while enabling sustainable innovation.
Business Impact
SRE transforms reliability into a growth driver. By focusing on critical customer journeys, it ensures that IT performance directly protects revenue and brand loyalty. The real impact is the ability to scale without proportional cost increases while accelerating time to market through safer, automated deployments. This shifts I&O from a “cost center” to a competitive advantage that manages risk at the speed of the business.
Drivers
  • Organizations are under pressure to meet customer requirements for reliability while scaling their digital services.
  • SRE has evolved since its initial implementations. Reliability remains the focus and integrates a rich body of knowledge that complements agile, DevOps and platform engineering approaches.
  • Organizations that have adopted highly skilled automation practices (usually DevOps) and usage of infrastructure-as-code capabilities to deliver digital business products expand to incorporate reliability into these products.
  • The most common use case, judging from inquiry calls with Gartner clients, is to leverage SRE concepts to improve the reliability of existing systems that are not meeting customer requirements for availability or performance, or proving difficult to scale.
  • As more autonomous agents are deployed, SRE provides the guardrails (SLOs) for nondeterministic AI behavior.
  • Moving from automation to autonomous operations uses agents to detect, diagnose and self-heal outages.
  • SRE is required to manage unpredictable dependencies created by multiagent architectures.
Obstacles
  • Belief that reliability is an I&O responsibility only, rather than being a shared organizational goal.
  • Organizations have difficulty shifting mindsets from service-level agreements with blanket availability statements to defining and measuring reliability in terms of customer-focused SLOs.
  • Defining appropriate SLOs that reflect both business needs and user expectations can be complex.
  • Finding SRE role candidates with the right mix of development, operations and people skills is a challenge for clients. This impacts initial adoption and scaling efforts.
  • Clients have voiced problems with product owners who overly focus on functional requirements, thus hampering improvements.
  • SRE requires close collaboration between development and operations teams. Organizational silos can make it difficult to implement SRE effectively.
  • SRE initiatives require executive sponsorship and investment to effectively scale.
User Recommendations
  • Define your goals for reliability that can guide implementation of SRE.
  • Begin your SRE journey with a small, focused initiative and then iteratively grow adoption.
  • Select opportunities that are politically friendly, will demonstrate sufficient value and have an acceptable risk profile.
  • Automate for SRE efficiency.
  • Work with the product owner to identify the SLIs and SLOs necessary to make customers and users happy.
  • Implement and improve observability to objectively report on performance relative to error budgets and SLOs.
  • Ensure product owners treat SLOs as a feature and are accountable for them.
  • Instill collaboration among site reliability engineers, developers and other stakeholders.
  • Create a community and leverage organizational learning practices while evolving SRE practices.
  • Cultivate a blameless culture where mistakes are seen as opportunities for learning and improvement.
  • Evaluate and invest in AI SRE tooling to lower cost of SRE adoption, meet operational demands and deliver effective reliability.
  • Establish “agentic guardrails” by defining SLOs specifically for AI agents to ensure autonomous actions don’t breach safety or cost boundaries.
Gartner Recommended Reading

Container Management

Analysis By: Dennis Smith
Benefit Rating: High
Market Penetration: 20% to 50% of target audience
Maturity: Mature mainstream
Definition:
Gartner defines container management as offerings that enable the development and operation of containerized workloads. Delivery methods include cloud, managed services, and software for containers running on-premises, in the public cloud and at the edge. Associated technologies include orchestration and scheduling, service discovery and registration, image registry, routing and networking, service catalog, management user interface, cluster fleet management, and APIs.
Why This Is Important
Container management automates the provisioning, operation, and life cycle management of containerized applications at scale. Centralized governance and security are used to manage container instances and associated resources. Container management supports the requirements of modern applications, including platform engineering, AI frameworks and associated workloads, cloud management and continuous integration/continuous delivery pipelines.
Business Impact
Demand for containers continues to rise, though the technology is maturing. Application developers and I&O teams practicing DevOps value containers for their ability to help agile application deployment by ensuring immutable and consistent environments throughout the development and delivery life cycle. As organizations make support for containers a standard part of their infrastructure, they require container orchestration platforms that can operate containers at scale, introducing the need for container management platforms. Container management is a foundational technology for AI workloads as they grow.
Drivers
  • The ongoing adoption of DevOps-based application development processes.
  • The growth of cloud-native application architecture based on microservices.
  • The increase of large-scale AI workloads requiring orchestration, for both training and inference.
  • System management approaches that are based on immutable infrastructure, which gives organizations the ability to update systems frequently and reliably rather than to repeatedly patch them.
  • Cloud-based services built with replaceable and horizontally scalable components.
  • The intersection of container technology and AI requirements, in which the scalability and elasticity provided by container management enables AI workload deployment.
  • A vibrant open-source ecosystem and competitive vendor market have culminated in a wide range of container management offerings.
  • Container-related edge computing use cases have increased in industries that need to get compute and data closer to the activity (such as telcos, retail stores, and manufacturing plants).
  • Cluster management tooling that enables the management of container nodes and clusters across different environments, including hybrid and multicloud scenarios, is increasingly in demand.
  • All major public cloud service providers now offer on-premises container solutions.
  • Independent software vendors (ISVs) increasingly package their COTS software for container management systems through container marketplaces.
Obstacles
  • More abstracted, serverless offerings may enable enterprises to forgo certain aspects of container management. These services embed container management in a manner that is transparent to the user.
  • Vulnerabilities throughout the software supply chain and dispersed responsibility across groups put container deployments at risk of security breaches.
  • Organizations that perform relatively little app development or make limited use of DevOps principles are better served by SaaS, ISV, or traditional application development packaging methods.
  • Additional cost and complexity in managing costs for public cloud container deployments is exacerbated by enterprises overprovisioning to ensure adequate compute and storage resources.
  • The trend of AI-driven software stack design, such as through vibe coding, which will likely lead to a further adoption of serverless/PaaS.
User Recommendations
  • Determine whether your organization is a good candidate for container management software adoption by weighing organizational goals of increased software velocity and immutable infrastructure.
  • Use container management capabilities integrated into cloud IaaS and PaaS providers’ service offerings by experimenting with process and workflow changes that accommodate the incorporation of containers.
  • Avoid using upstream orchestrators directly (such as self-developed Kubernetes platforms without vendor support) unless the organization has adequate in-house expertise to support it.
  • Anticipate the need for cost optimization tooling and techniques for deployments in the public cloud.
  • Invest in cloud-native specialized observability tools and services for resource management and troubleshooting.
  • Use DevSecOps practices and security guardrails to identify and address vulnerabilities in the container-deployment process.
Sample Vendors
Alibaba Cloud; Amazon Web Services; Broadcom (VMware); Google; IBM; Microsoft; Mirantis; Oracle; Red Hat; SUSE
Gartner Recommended Reading

Observability Platforms

Analysis By: Padraig Byrne, Gregg Siegfried
Benefit Rating: Transformational
Market Penetration: 20% to 50% of target audience
Maturity: Early mainstream
Definition:
Observability platforms are products used to understand the health, performance and behavior of applications, services and infrastructure. They do this by ingesting telemetry from a variety of sources, including logs, metrics, events and traces. Observability platforms enable analysis of the ingested telemetry, either via human operators or machine intelligence, to determine changes in system behavior that impact end-user experience, such as outages or performance degradation.
Why This Is Important
Modern software environments, spanning cloud-native applications, AI workloads and distributed systems, generate complex, dynamic behavior that legacy monitoring tools cannot explain. Observability platforms give engineering and operations teams the ability to ask and answer novel questions about system behavior, accelerating incident resolution, supporting developer self-service and enabling organizations to operate AI-driven services with the same confidence as traditional applications.
Business Impact
Observability platforms reduce the frequency and severity of service outages, while improving software quality by surfacing previously invisible defects before service degradation occurs. As AI-assisted workloads and distributed architectures increase system complexity, the business value of rapid, evidence-based incident resolution continues to grow.
Drivers
  • The complexity of modern distributed systems (microservices, serverless, Kubernetes and multicloud environments) makes traditional threshold-based monitoring insufficient. Observability platforms enable engineering teams to investigate novel failure modes without requiring prior knowledge of what to instrument.
  • OpenTelemetry has become the de facto standard for instrumenting cloud-native and AI applications, with all major observability platform vendors now supporting OpenTelemetry-native ingestion. OpenTelemetry’s scope now extends to continuous profiling and AI/LLM telemetry, consolidating its role as the single instrumentation framework across traditional and emerging workload types.
  • AI applications and agentic systems introduce new observability requirements. Organizations deploying autonomous agents need to observe model behavior, tool invocations and reasoning chains. This demand is driving both new platform capabilities and new vendor categories.
  • Platform engineers are centralizing observability as a shared internal service. I&O teams are increasingly responsible for providing observability tooling as part of an internal developer platform, driving standardization and consolidation across previously fragmented tool estates.
  • Telemetry data volumes from cloud-native and AI workloads have grown faster than budgets, making cost management a primary buying criterion. Capabilities such as adaptive sampling, pipeline filtering and tiered storage are now as important as analytical capabilities.
Obstacles
  • Telemetry data volumes from cloud-native and AI workloads continue to outpace budgets, giving rise to the “observability tax”: the growing share of engineering spend consumed by data ingestion, storage and retention costs. Many organizations are adopting observability pipeline tools as a secondary investment to manage this burden, adding unplanned complexity to their tool estates.
  • Observability tooling stickiness remains high. Significant investment in existing platforms, combined with deeply embedded instrumentation, creates organizational inertia that slows migration even when better options exist.
  • Observing AI applications requires skills most ITOps and SRE teams do not yet possess. Understanding nondeterministic model behavior, evaluating trace data from agentic workflows and interpreting LLM-specific signals demand expertise that is scarce and unevenly distributed.
  • Market consolidation through acquisition has introduced uncertainty. Several independent observability vendors have been absorbed into larger portfolios, raising concerns about roadmap continuity, pricing changes and long-term platform viability.
User Recommendations
  • Prioritize vendors that support OpenTelemetry natively for collection and instrumentation, reducing lock-in risk and improving portability across a consolidating vendor landscape.
  • Assess platforms’ built-in cost management capabilities, including adaptive sampling, tiered storage and pipeline filtering, before assuming that ingestion-based cost growth is unavoidable.
  • Tie service-level objectives to desired business outcomes using specific metrics, and use observability platforms to understand and communicate differences to business leaders.
  • Use complementary approaches, such as digital experience monitoring and synthetic monitoring to close visibility gaps for cloud, especially SaaS services.
  • Treat observability as a shared platform engineering capability rather than a per-team tooling decision, driving standardization and reducing the operational overhead of managing fragmented tool estates.
Sample Vendors
Coralogix; Datadog; Dynatrace; Grafana Labs; Honeycomb; New Relic; Palo Alto Networks (Chronosphere)
Gartner Recommended Reading

Secrets Management Tools

Analysis By: Paul Mezzera, Steve Wessels
Benefit Rating: Moderate
Market Penetration: 20% to 50% of target audience
Maturity: Adolescent
Definition:
Secrets management tools programmatically issue, store, retrieve, rotate and manage secrets like keys, passwords, OAuth client credentials and certificates. They manage these secrets for workloads such as containers, applications, services, scripts, DevOps pipelines and AI agents. These tools provide service through APIs, command line interfaces (CLIs) and software development kits (SDKs). They are delivered as software or as a service and include a secure and encrypted vault.
Why This Is Important
Machine-to-machine communication is ubiquitous, with various workloads like containers, applications, services, functions and AI agents requiring access to sensitive data and infrastructure. Failing to secure this access can lead to — and has led to — security breaches. Workloads use credentials that must be protected to prevent exposure. Secrets management tools are essential in environments where API keys and similar secrets are prevalent, securely issuing and storing them.
Business Impact
Effective secrets management enhances security for transactions, operations and communications between nonhuman actors, such as workloads. It mitigates risks of breaches by safeguarding and rotating secrets used to implement controls and access restrictions. It also manages IAM technical debt, serving as a pragmatic necessity for workloads and DevOps pipelines where the use of API keys and similar secrets remains prevalent.
Drivers
  • Threat landscape: Secrets used by workloads often contain privileged entitlements and grant access to critical systems and sensitive data. Unprotected secrets are a lucrative target for cybercriminals. Securing secrets is now a priority after years of neglect and resulting secrets sprawl.
  • Operational efficiency: Managing secrets using manual processes, such as storing them in configuration files or using spreadsheets or workforce password managers, is both inefficient and insecure. Secrets management tools automate issuing, storing and retrieving secrets, improving operational efficiency by automating and reducing the manual administrative burden.
  • Cloud and API adoption: Continuing migration to cloud and hybrid environments requires secrets to be managed across distributed environments, which entails consistent visibility and control over secrets across cloud infrastructure providers, container orchestration systems, APIs and on-premises systems.
  • AI and AI agents: The rise of AI and AI agents, which require access to sensitive data and systems, further underscores the need for effective secrets management to safeguard credentials used by these applications.
  • Regulatory landscape: Many regulatory frameworks mandate protection of sensitive data. When workloads access sensitive data, credentials must be protected. Regulations also require management of cryptographic keys, which secrets management tools can secure.
  • DevOps: The adoption of DevOps requires developers to securely access secrets during build, deploy and runtime. Secrets management tools eliminate storing secrets in configuration files or code repositories and integrate with CI/CD pipelines to automate build, test, delivery and deployment.
  • Auditability and accountability: Secrets management tools provide audit logs and monitoring to track and monitor the use of secrets. This enhances accountability, supports forensic investigations and enables compliance audits with clear records of secrets-related activities.
Obstacles
  • Lack of visibility and observability: Organizations struggle to track all credentials and secrets used across their infrastructure due to insufficient discovery and continuous monitoring mechanisms.
  • Vendor lock-in and interoperability: Proprietary APIs and SDKs make it difficult to switch vendors, leading to dependency issues.
  • Fragmentation: Organizations often use multiple secrets management tools (commonly provided by cloud service providers [CSPs]), sometimes chosen independently by teams, leading to inconsistent policies and the need for standardization.
  • Overreliance on secrets management: Relying solely on secrets management can cause vendor lock-in and hinder unified machine identity strategies, often resulting in static, hard-to-manage credentials instead of adopting broader identity management solutions.
  • Pricing: Inconsistent pricing models, especially per-client fees for stand-alone tools, make adoption costly and complicate budgeting compared to native CSP solutions.
User Recommendations
  • Eliminate reliance on static secrets, especially in new deployments: Minimize the use of static secrets where possible and prioritize solutions that support dynamic secret generation and automated rotation to enhance workload-to-workload security.
  • Embrace fragmentation: Don’t rely on a single secrets manager, which can lead to vendor lock-in and poor fit for diverse needs. Instead, prioritize features that provide observability and governance and platform-managed workload identities across multiple tools, vaults and clouds. Begin by discovering secrets across platforms; then, expand to cover exposed locations like cloud apps, collaboration tools and chats.
  • Prohibit storing credentials in cleartext: Workloads need to identify themselves to the secrets management tool, often through an authentication mechanism. Do not authenticate a workload to a secrets manager using a credential stored at rest within the workload.
Sample Vendors
Akeyless; Amazon Web Services; ARCON; Delinea; Google; IBM (HashiCorp); Infisical; Microsoft; Palo Alto Networks (CyberArk)
Gartner Recommended Reading

Chaos Engineering

Analysis By: Jim Scheibmeir, Hassan Ennaciri
Benefit Rating: Moderate
Market Penetration: 20% to 50% of target audience
Maturity: Adolescent
Definition:
Chaos engineering (CE) is the use of experimental and potentially destructive failure testing or fault injection to uncover vulnerabilities and weaknesses within a distributed system. Chaos engineering tools provide the ability to systematically plan, document, execute, and analyze an attack on components and whole systems throughout a system’s life cycle.
Why This Is Important
Many organizations rely on test plans that overemphasize functionality and underemphasize validating the system’s reliability and resilience. The distribution and complexity of systems make understanding them more difficult. CE shifts the focus of testing a system from the “happy path” toward testing it under “chaotic path” conditions by intentionally simulating failures. Proactive CE identifies potential system improvements for confidentiality, integrity, and availability.
Business Impact
CE is aimed at minimizing time to recovery and the change failure rate, while maximizing uptime and responsiveness. Addressing these elements helps improve customer experience, satisfaction, retention, and acquisition. Improving systems reliability also helps traditional cybersecurity concerns of confidentiality, integrity, and availability.
Drivers
  • As applications become “intelligent by design,” the need for making them “resilient by design” increases. Failure injection capabilities for large language models (LLMs) and AI agents will be the next driver for this practice and its associated tools and vendors.
  • Increased complexity of systems and increasing customer expectations are the two largest drivers of CE and the associated tools.
  • As systems become richer in features, they also become more complex in their composition and more critical to digital business success.
  • Overall, CE enhances organizational resilience by improving the way processes, knowledge, and technology are managed and continuously adapted.
  • Teams often lack the confidence to handle failures and the psychological safety to take action to resolve incidents. CE can help build that confidence.
  • More resilient systems allow support and development teams a better work-life balance, less unplanned work, and more consistency in their ability to deliver on planned work.
Obstacles
  • Within many organizations, the predominant view of CE is that the practice is random, first implemented during production, and increases, rather than reduces, risk.
  • Organizational culture and attitudes toward quality and testing can present barriers to adopting CE. When quality and testing are only viewed as overhead costs, there will be a focus on feature development over application reliability.
  • It can be challenging to secure the time and budget to invest in learning CE and associated technologies. Organizations must reach minimum levels of expertise so that value is returned.
  • There are costs associated with CE and system reliability that can’t be ignored. Not every process in the system demands the same level of resiliency; the focus should be on processes that are most integral to the needs of the business.
User Recommendations
  • Utilize a test-environment-first approach by practicing CE in preproduction environments.
  • Incorporate CE into your system development, continuous integration/continuous delivery, or testing processes.
  • Leverage CE when embedding generative AI API calls in your applications to test fallback patterns.
  • Implement CE to prepare your organization against ransomware-style attacks.
  • Utilize scenario-based tests — known as “game days” — to evaluate and learn how individual IT systems would respond to certain types of outages, including catastrophic failures.
  • Prioritize CE activities on critical systems that have elevated security privileges, business-critical services such as payment/payroll, or components that are single points of failure.
  • Investigate opportunities to use CE in production to facilitate learning and improvement at scale as the practice matures.
  • Adopt a platform or tool to track activities and create metrics to build feedback for continuous improvements.
Sample Vendors
Amazon Web Services; Gremlin; Harness; Microsoft; Quinnox; Steadybit
Gartner Recommended Reading

Internal Developer Portals

Analysis By: Cary Pillers
Benefit Rating: High
Market Penetration: More than 50% of target audience
Maturity: Mature mainstream
Definition:
Internal developer portals (IDPs) are the front end for internal developer platforms. They provide developers with self-service capabilities to discover, access and operate platform tools, services and knowledge. IDPs improve developer experience and service reliability while enabling centralized governance and a common, up-to-date view of service ownership, dependencies and standards. Capabilities include software catalogs, software quality, scorecards and scaffolding templates.
Why This Is Important
IDPs help developers navigate complex development environments, understand service and organizational interdependencies across software systems, and streamline software delivery through:
  • Common catalogs and information shared across development teams.
  • Self-service access to and automation of platform components and environments.
  • Centralized tracking of progress against reliability, compliance and security requirements.
  • Context and consistency for AI-native tools and automation.
Business Impact
  • Developer experience and productivity — Helps developers improve their delivery cadence by improving developer experience, reducing cognitive load and shortening the feedback loop.
  • Reliability and resilience — Aims to provide visibility to application health and includes scorecards to enforce the production readiness standards.
  • Security and governance — Includes prebuilt toolkits, templates and curated libraries that help create “paved roads” with built-in compliance, security and audit policies.
Drivers
  • Platform engineering — Organizations are adopting platform engineering principles and creating platform teams to scale cross-cutting capabilities for multiple development teams. Platform teams curate internal developer platforms to rationalize and unify the siloed systems and processes. IDPs serve as the user interface through which developers can utilize the capabilities of internal developer platforms.
  • Backstage Backstage is one of the first open-source frameworks for building developer portals. Created at Spotify and now a Cloud Native Computing Foundation (CNCF) project, it has played a significant role in shaping how organizations approach IDPs. The thriving open-source community supporting Backstage has largely contributed to its enormous mind share and rapid adoption. This is despite the effort required for initial deployment and ongoing maintenance of Backstage.
  • Developer experience — A great developer experience that accelerates software development is a key competitive advantage. Software engineering leaders are increasingly focused on minimizing developer friction and frustration. The ability to curate and provide customizable, developer-friendly experiences within the developer portal and rein in complexity helps reduce the cognitive load for developers.
  • Self-service — Enables developers to create, modify, and deploy services without the need for tickets or human interaction. This enables automated provisioning of resources safely, consistently and repeatably.
  • AI agents as platform consumers — As AI coding and operational agents proliferate, IDPs provide a consistent source of guidance, standards, and contextual information that both developers and AI agents must follow and reuse. As AI agents increasingly operate as platform consumers, IDPs can support discovery of available agents and provide lightweight user experiences to manage agent instances and tasks, or link to underlying platform workflows at the team or developer level.
Obstacles
  • Failure to be developer-centric — Organizations that don’t take developer pain points and strategic business goals into account while determining portal capabilities fail to see a return on their investment.
  • Required integrations are often ignored — Portals that don’t integrate with the software development life cycle tools for building, deploying and operating software struggle to increase adoption and be relevant.
  • Platform teams must be established — A dedicated platform team led by a platform product owner is needed to manage and evolve the portal and ensure it meets desired objectives. Without a dedicated owner, gaps emerge between developer expectations and the portal’s capabilities.
  • Portals are not free — Whether using a framework or solution from a vendor, organizations trade build, maintenance, and administration costs for configuration, subscription, and administration costs. The choice must be based on available resources, their skills and time-to-market priorities.
User Recommendations
  • Reduce the risk of poor IDP adoption by applying product thinking, ensuring the product owner for the platform works closely with developers to understand their ways of working and integrate IDP capabilities into existing workflows.
  • Evaluate commercial IDPs by assessing their pros and cons against these criteria: flexibility and open ecosystem, integrated portal capabilities as part of a broader developer platform, out-of-the-box capabilities without additional customization or development work.
  • Go beyond developer roles to serve multiple personas and teams involved in delivering technology solutions by enabling discovery and access to underlying platform capabilities that streamline software engineering, data engineering application security and IT operations workflows.
  • Consider the risk that generative AI may transform user interfaces, and the impact that may have on IDP functionality.
Sample Vendors
Cortex; DataDog; Harness; OpsLevel; Port; Red Hat; Roadie; Spotify; WSO2
Gartner Recommended Reading

Threat Modeling Automation

Analysis By: William Dupre
Benefit Rating: High
Market Penetration: 20% to 50% of target audience
Maturity: Early mainstream
Definition:
Threat modeling automation tools help with the creation of security requirements and threat models. Such tools highlight potential security ramifications of application architectures and recommend secure coding practices or architectural mitigations. They also can manage threat libraries, track mitigations and integrate security analysis into the software development life cycle (SDLC).
Why This Is Important
Threat modeling is key to creating applications that are secure by design. Automated tools enhance the threat modeling process in the design phase by accelerating threat identification and improving consistency. Although they do not secure applications, automated tools help in the creation of secure application architectures and ensure identification of appropriate and specific security requirements.
Business Impact
Threat modeling automation tools significantly decrease the effort required to create and maintain threat models, security requirements and risk assessments. This ensures early definition of security requirements that are specific to individual projects while costs and risks are low, rather than later in the development process. This approach offers significant benefits to multiple groups within an organization, including architects, developers, security teams and even business stakeholders.
Drivers
  • Organizations of all kinds continue to struggle to create secure applications. Issues include inadequate security capabilities, such as authentication, access control and data protection, and fundamental security design flaws, all of which leave applications vulnerable to attack.
  • Secure-by-design initiatives and secure software compliance mandates make threat modeling an essential practice. Automated threat modeling tools make the practice faster and more scalable. They also help ensure that specific requirements associated with mandates and regulations are addressed.
  • Modern applications — incorporating distributed cloud-native technologies, increased use of internal and third-party APIs, and agentic AI — are more complex and, as a result, prove difficult to manually and accurately model for threats. Threat modeling automation speeds the threat modeling process and helps modelers identify threats and relevant countermeasures.
  • The rapid pace of development limits the time and resources available for threat modeling. Iterative approaches to development, such as agile practices, mean that threat models must be updated more frequently. Both factors strain manual approaches and increase the likelihood that threat modeling will be limited in scope or simply skipped entirely.
  • Vendors are adding AI capabilities into threat modeling and other application security products to enable better automation and identification of threats in the software development process.
Obstacles
  • The ability to accurately represent a rapidly changing application remains a weak spot of threat modeling automation tools. Most of today’s tools require user intervention to update models as applications change, which leads to abandonment. This is improving as vendors begin to link systems directly to cloud platforms or infrastructure-as-code files, ensuring that changes are reflected automatically in the model, which will then automatically produce updated guidance.
  • Capabilities vary. Free and open-source tools enable easy adoption but fall short when modeling more complex systems.
  • Most organizations still focus on application security testing as they establish an application security program. These tools are essential to identifying vulnerabilities in code during the development stage but fail to identify design flaws during planning.
User Recommendations
  • Treat threat modeling and security requirement generation as primary practices within a secure SDLC. Threat modeling automation tools can help accelerate these tasks while incorporating threat and security requirement knowledge from tool vendors.
  • Use these tools to automate manual or overlooked efforts. Doing this ensures that threat modeling and security requirement generation activities are incorporated into the development workflow. Test cases must then be created to ensure security requirements are effectively covered.
  • Evaluate emerging AI-based capabilities that support threat modeling. AI offers the potential for improved analysis and efficiency, threat identification and for simplified interaction with the tools.
  • Train development, engineering, operations and architectural staff in the use and value of threat modeling automation tools. Encourage their use early and continuously in the development process and after deployment to validate application threat protection efforts.
  • Define success metrics and ROI upfront. Establish clear goals for threat identification, process improvement and reporting before implementation to align on value measurement and secure ongoing executive support.
Sample Vendors
Aristiun; SecureFlag; Security Compass; ThreatModeler; TrustOnCloud; Tutamantic Sec
Gartner Recommended Reading

Service Mesh

Analysis By: Shameen Pillai
Benefit Rating: Low
Market Penetration: 1% to 5% of target audience
Maturity: Adolescent
Definition:
A service mesh is a distributed computing middleware that manages communications between application services — typically within managed container systems. It provides lightweight mediation for service-to-service communications and supports functions such as authentication, authorization, encryption, service discovery, request routing, load balancing, self-healing recovery and service instrumentation.
Why This Is Important
A service mesh is middleware for managing and securing service-to-service (east-west) communications, particularly in microservices and Kubernetes environments. It provides deep visibility into service interactions, enabling proactive monitoring, rapid diagnostics and enhanced resilience. By automating complex communication tasks — such as authentication, encryption and traffic management — a service mesh boosts developer productivity and ensures consistent enforcement of security and operational policies across distributed applications. As organizations adopt Zero-Trust security models and scale across multiple clusters, service mesh plays a critical role in simplifying operations, reducing risk and unifying governance in modern, cloud-native architectures.
Business Impact
By automating service discovery, security, traffic management, and observability, service meshes provide the robust infrastructure needed for reliable “Day 2” operations — ensuring stability, compliance and resilience as microservices scale. This enables organizations to confidently operate complex, distributed systems and adopt Zero-Trust security models. However, due to its complexity and operational overhead, service mesh is best suited for larger, dynamic deployments; for smaller or less complex environments, it may be unnecessary.
Drivers
  • Microservices and containers: Service mesh adoption is driven by the rise of microservices architectures and container orchestration platforms like Kubernetes. Service meshes enable dynamic service discovery, secure inter-service communication with mTLS, and robust management in ephemeral, highly dynamic environments.
  • Observability: As microservices scale, DevOps teams require advanced observability to monitor operations, trace errors, and anticipate issues. Service meshes provide built-in instrumentation, delivering logs and metrics to centralized dashboards for proactive management.
  • Resilience: Service meshes implement critical communication stability patterns — such as retries, circuit breakers, and bulkheads — supporting self-healing, canary and blue-green deployments to enhance application reliability.
  • Bundled feature: Many modern container management systems and hyperscale cloud providers now bundle service mesh capabilities, encouraging adoption by integrating them with other cloud-native services.
  • Federation and multienvironment support: Independent vendors (e.g., Buoyant, Greymatter.io, IBM, Kong, Solo.io) offer service meshes that support multi-cluster and multi-environment deployments, enabling consistent policy enforcement and management across diverse infrastructures.
Obstacles
  • Not always necessary: While valuable for complex microservices in Kubernetes, service mesh is not required for all deployments and may be excessive for smaller environments.
  • Complexity: Service mesh introduces significant operational and administrative complexity. This has led to growing debate within the tech community about its necessity and the overhead it brings. Complexity is particularly a concern for multienvironment service mesh.
  • Redundant functionality: Overlap with ingress controllers, API gateways and other proxies causes confusion. Interoperability and unified management across these layers remain immature in the vendor ecosystem.
  • Resource overhead: Traditional sidecar-based service meshes add resource and latency overhead. Newer architectures (e.g., eBPF, ambient mesh, shared-agent models) aim to reduce this, but may compromise observability or introduce new trade-offs.
  • Vendor competition: Independent service mesh solutions face stiff competition from platform-integrated options bundled by major cloud and container providers, which can limit adoption and market differentiation.
User Recommendations
  • Carefully assess whether the benefits of improved security, observability and resilience from a service mesh justify the added complexity and operational overhead. Service mesh value increases with the scale and complexity of east-west service interactions.
  • Prefer service meshes that are natively integrated with your container management platform, unless you require advanced federation or multienvironment support.
  • Assign service mesh ownership to a cross-functional platform engineering team that collaborates closely with networking, security and development groups to reduce friction and ensure alignment.
  • Foster knowledge sharing and consistent security policy enforcement by working with infrastructure, operations and security teams managing related technologies like API gateways and application delivery controllers.
  • Regularly evaluate emerging service mesh architectures (e.g., sidecarless/ambient mesh) to balance resource efficiency with observability and operational needs.
Sample Vendors
Amazon; Buoyant; Cisco (Isovalent); CNCF (Istio); Google; IBM (HashiCorp); Istio; Kong; Microsoft; Solo.io
Gartner Recommended Reading

Climbing the Slope

Cloud-Native Application Protection Platforms

Analysis By: Dale Koeppen, Neil MacDonald
Benefit Rating: High
Market Penetration: 20% to 50% of target audience
Maturity: Early mainstream
Definition:
Cloud-native application protection platforms (CNAPP) are an integrated set of security and compliance capabilities designed to help secure and protect cloud-native infrastructure and applications across development and production. CNAPP consolidates a broad range of security and compliance functions, extending protection from DevOps pipelines and code development through to cloud infrastructure and workload runtime environments.
Why This Is Important
Securing cloud environments often requires managing poorly integrated tools from multiple vendors, leading to development friction, fragmented risk visibility and alert fatigue. CNAPP provides a single integrated solution to protect the cloud application life cycle and deliver a unified, prioritized view of risk across cloud-native application pipelines and operational workloads.
Business Impact
CNAPPs consolidate disparate, fragmented security scanning and protection tools that drive up cost and complexity for IT teams and impede remediation workflows. Using a CNAPP offering improves developer and security professional efficacy. It also reduces complexity, streamline communication and gain better visibility into cloud risks, while maintaining development agility and improving the developer’s experience.
Drivers
Cloud-native application protection platforms provide an integrated set of cloud risk management capabilities that:
  • Reduce the chance of misconfiguration, mistake or mismanagement as cloud-native applications are rapidly developed, released into production and iterated.
  • Consolidate and simplify security tools and vendors throughout the software development life cycle and workload runtime analysis.
  • Reduce the complexity and costs associated with creating secure and compliant cloud-native applications.
  • Facilitate automated and continuous reporting and auditing of cloud security posture and status.
  • Improve developer acceptance with security-scanning capabilities that seamlessly integrate into their development pipelines and tooling.
  • Emphasize scanning proactively in development and rely less on runtime protection, which is well-suited for container as a service and serverless function environments.
  • Streamline communications to improve remediation efforts and reduce operational risk.
  • Address security posture management of emerging AI services offered by cloud providers.
Obstacles
  • Although runtime-focused vendors are mature in workload protection, they aren’t necessarily good at integrating into DevOps.
  • Vendors focusing on DevOps and pipeline security are not necessarily mature in providing workload runtime protection.
  • Containers and serverless functions typically benefit from adaptable security options rather than relying on heavyweight runtime protection.
  • No single CNAPP offering does everything. Convergence of capabilities will occur but will take place over several years.
  • Organizations often have different teams independently selecting application security and cloud posture tools, while separate teams manage the runtime and protection aspects of workloads.
  • Buying centers and influencers are shifting to newer roles, such as DevOps architects, cloud security architects, AI strategy officers and platform engineering, requiring information security teams to coordinate with different stakeholders.
User Recommendations
  • Prioritize vendors that meet essential CNAPP criteria by providing a balance of proactive security capabilities for DevOps and adaptable, reactive runtime security options for SecOps.
  • Select integrated offerings with flexible licensing models that allow you to pay only for the capabilities your organization is prepared to use.
  • Evaluate whether the vendor can provide customized alerting, reporting and recommendations suited to the needs of stakeholders across Development, Security and Operations (DevSecOps).
  • Consolidate open-source vulnerability scanning and software composition analysis by integrating or replacing these tools within your CNAPP solution.
  • Proactively scan containers in development pipelines for all types of vulnerabilities, not just vulnerable components, including hard-coded secrets, malware and Kubernetes misconfiguration.
  • Choose CNAPP providers that, despite gaps, have strong integrations with complementary security tools to ensure seamless risk context sharing and consolidation.
Sample Vendors
Aqua Security; CrowdStrike; Fortinet; Orca Security; Palo Alto Networks; Qualys; Sysdig; Tenable; Trend Micro; Wiz
Gartner Recommended Reading

Open-Source Program Office

Analysis By: Nitish Tyagi, Arun Chandrasekaran, Mark O'Neill
Benefit Rating: Moderate
Market Penetration: 5% to 20% of target audience
Maturity: Adolescent
Definition:
An open-source program office (OSPO) directs the strategies for governing, assessing, managing, promoting and efficiently using open-source software (OSS) and open-source data, standards or models. The OSPO is led by a program leader who typically reports to the CTO, CIO or software engineering leader, and includes members from legal, security, software engineering or application delivery, and procurement.
Why This Is Important
An OSPO governs OSS and open generative AI (GenAI) model use, mitigates legal and security risks, enables strategic adoption, boosts developer productivity and ensures compliance. It acts as a central hub for open-source expertise, facilitating both consumption and contribution. Ultimately, it helps organizations confidently leverage OSS and open GenAI models. As sovereignty concerns drive organizations to embrace open-source software more strongly, OSPOs become even more important.
Business Impact
OSS enables innovation, accelerated software development, cost savings and talent retention. An OSPO ensures efficient OSS consumption and contribution, implements governance and promotes the value of OSS to stakeholders. Also, an OSPO helps to develop innersource practices, and safe and efficient adoption of open GenAI models. As the OSPO matures, it becomes a strategic partner in all technology decisions.
Drivers
  • Management of OSS usage and risks: OSS is widely used, often without full awareness. An OSPO provides a central point for building strategies to govern this use, mitigate security vulnerabilities, ensure license compliance and address legal risks of OSS.
  • Increased adoption of open GenAI models: Open GenAI or open-weight models offer cost-efficiency and higher control on inferencing, but come with compliance and security nuances. In the AI-native era, OSPOs are responsible for OSS and open GenAI.
  • Strategic alignment: An OSPO aligns OSS strategy with the overall business and IT strategies. It articulates the role of OSS in digitalization and ensures that its use supports business goals.
  • Fostering innovation: OSS is the backbone of digital innovation. An OSPO helps organizations strategically leverage OSS to drive innovation and adopt cutting-edge technologies like AI, cloud and DevOps.
  • Developer productivity and talent: An OSPO streamlines OSS consumption, reducing the burden on development teams. It also attracts and retains talent by offering opportunities to work on cutting-edge OSS projects and contribute to communities.
  • Cost-efficiency: Leveraging OSS can lead to cost savings, but effective governance is key. An OSPO helps optimize the TCO associated with OSS.
  • Innersourcing and collaboration: OSPOs often drive innersourcing, applying OSS principles internally to increase code reuse, knowledge sharing and collaboration across teams.
  • Meeting regulatory and customer requirements: With increasing focus on software supply chain security, OSPOs can help organizations generate software bills of materials (SBOMs) and meet disclosure requirements from customers and regulatory bodies.
  • EU digital sovereignty initiative: Increased focus on digital sovereignty is leading to increased adoption of open source in the European region, resulting in higher importance of OSPOs. Additionally, the UN’s recent initiative, OSPO for Good, is increasing awareness of OSPOs.
Obstacles
  • Lack of funding and executive sponsorship are the biggest reasons for the limited adoption of OSPOs. Making the case for funding an OSPO is difficult because an OSPO does not directly create software and, therefore, may be seen as only a cost center.
  • Measuring the value of an OSPO in a shorter period of time (zero to two years) is challenging because of the delay in metrics collection. This also reduces the leadership’s confidence.
  • Siloed organizations and teams make it difficult for an OSPO to drive collaboration.
  • Getting teams on board with cultural change may be a challenge, as OSS governance could be interpreted as a means to restrict software development freedom.
  • Focusing on OSS consumption either decentralizes the operations after a certain time or sets up a subcommittee, rather than a fully fledged OSPO.
  • Navigating open-source licensing models is challenging without expertise. This complexity increases with the rise of source-available licenses.
User Recommendations
  • Treat OSS as a strategic investment by establishing an OSPO for building policies and processes to govern, assess and manage OSS.
  • Allocate open-source champions to communicate OSPO’s messaging across the organization and relay end-user needs back to the OSPO.
  • Use OSPO’s capabilities with safe and efficient adoption of open GenAI models and open-weight models such as DeepSeek, OSS-GPT and Qwen.
  • Define the correct set of metrics to measure OSPO success, such as OSS and innersource adoption, upstream code contributions, and hire and retention rates. Formulate and enforce a governance policy for consumption, contribution and OSS creation by building an OSS governance committee.
  • Work with platform engineering teams to provide the correct set of tools for artifact repository, security, issue tracking, continuous integration, continuous development, collaboration and knowledge management.
  • Evaluate your OSPO’s maturity level and execute strategies to gradually promote it to the next level.
Sample Vendors
The Apache Software Foundation (ASF); Bitergia; Cloud Native Computing Foundation (CNCF); The Linux Foundation; TODO Group
Gartner Recommended Reading

Team Topologies

Analysis By: Peter Hyde, Nabeeha Ahmed
Benefit Rating: High
Market Penetration: 20% to 50% of target audience
Maturity: Early mainstream
Definition:
Team Topologies originated in 2019 as an adaptive, team-first approach to creating high-quality products and services. Four team types and three interaction modes combine with Cognitive Load Theory and Conway’s Law to create an effective model for organizing teams that develop and operate software systems. Matthew Skelton and Manuel Pais created and popularized these concepts in their book “Team Topologies: Organizing Business and Technology for Fast Flow of Value.”
Why This Is Important
Autonomous agile teams with clear ownership of the products or services they deliver achieve higher quality and productivity than traditional teams. Scaling these teams to an empowered, efficient and effective product structure while retaining agility and innovation remains a significant challenge for engineering leadership. The Team Topologies approach provides a shared organizational design language for software engineering to achieve fit-for-purpose structures that fulfill business demand.
Business Impact
The Team Topologies approach supports the evolution of team structures and interactions within software engineering to guide organizational growth. It aims to achieve better alignment, increased productivity and improved business agility. Moving to a value-stream-aligned model with self-service platforms and empowered enablement teams helps improve developer experience, evolve engineering practices, increase business collaboration, build a better culture and deliver an aligned strategic vision.
Drivers
  • Reducing cognitive load: The adoption of modern, distributed architectural patterns and software delivery practices means that the process of developing and delivering software involves more tools, subsystems and moving parts than ever before. This cognitive load places a burden on product teams to build a delivery system in addition to the actual software they are trying to produce. The Team Topologies approach advocates for platform engineering to reduce the burden of infrastructure construction and maintenance.
  • Improving system architecture: The mirroring of an organization’s structure to its product architecture, as specified in Conway’s Law, is significant, as it enables an intentional evolution through organizational design to deliver the target architecture. This practice is referred to as the “Inverse Conway Maneuver.” It creates the impetus for architectural change by breaking down the silos that constrain the ability of individual teams to create value.
  • Scaling ways of working: Agile transformation, project-to-product transition and the need to successfully scale software engineering are introducing additional complexity. This complexity produces friction that slows the delivery of business value and reduces customer impact. Better product and service delivery through the Team Topologies approach creates aligned value streams that can iteratively and incrementally deliver better results.
Obstacles
  • Existing power structures: Traditional static hierarchies may delay or even block organizational transformation due to a fear of change and a perceived reduction in authority.
  • Lack of organizational design experience: Many software engineering leaders and coaches lack awareness of modern organizational models and do not have practical experience in organizational design change initiatives.
  • Complex legacy architectures: The challenge of establishing clear boundaries within monolithic architectures, which are often burdened by technical debt and manifold dependencies, hinders the formation of autonomous teams accountable for independent systems.
  • Sponsorship and funding: Gaining the authority, sponsorship and funding required to enact enterprise-level change is challenging. Incremental evolution through the Team Topologies approach lowers the entry bar but still impacts delivery speed during adoption.
User Recommendations
  • Form stream-aligned teams that are accountable for a single valuable product or service. Stream-aligned teams deliver rapid customer value securely and independently, with fast end-user feedback.
  • Reduce the cognitive load for stream-aligned teams by extracting supporting services into thinnest viable platforms. Platform teams empower stream-aligned teams to focus on creating stakeholder value.
  • Accelerate the flow of business change by identifying capability gaps and forming enabling teams to address them. Enabling teams provide training, coaching and mentoring to help other teams improve and become autonomous.
  • Review your organizational structure quarterly with Team Topology concepts to assess its effectiveness and ensure it remains fit for purpose. Monitoring cognitive load and interteam friction enables teams to optimize the flow of business value.
Gartner Recommended Reading

Infrastructure Automation

Analysis By: Chris Saunderson
Benefit Rating: High
Market Penetration: 20% to 50% of target audience
Maturity: Mature mainstream
Definition:
Infrastructure automation enables DevOps, platform, and infrastructure and operations (I&O) teams to deliver automated infrastructure services across on-premises and cloud environments. This includes the definition of environments, life cycle of services through creation, configuration, operation and retirement. These services are then made available through automation workflows, specialized platforms, self-service catalogs, direct invocation, agentic AI and API integrations.
Why This Is Important
Infrastructure automation delivers velocity, quality, and efficiency through scalable, declarative approaches for deploying and managing infrastructure. These tools integrate into delivery pipelines and platforms targeting deployments that range from on-premises, hybrid and to the cloud, enabling infrastructure consumers to build what is needed when needed. Once deployed, IA provides day-two-and-beyond operational automation, and extends to policy compliance and policy enforcement capabilities.
Business Impact
IA services enable:
  • Agility, scalability and collaboration: Fostering team autonomy, delivering what product teams need with security, scalability, cost and compliance requirements
  • Productivity: Version-controlled, declarative, repeatable, and efficient deployments
  • Cost improvement: Reductions in manual efforts through increased automation
  • Risk mitigation: Compliance driven by standardized configurations and governance
  • Quality assurance: Removal of human error with automated repeatable deployments
Drivers
IA tools support maturation beyond simple deployments through:
  • Hybrid infrastructure delivery that is vendor- and platform-neutral
  • Support for immutable and programmable infrastructures
  • Predictable delivery, enabling automated operations
  • Self-service and on-demand environment creation
  • Integration into DevOps, platform and infrastructure platform initiatives
  • Resource provisioning, including scale-up and scale-down cost optimization capabilities
  • Operational configuration management efficiencies
  • Policy-based delivery and assessment or enforcement of deployments against internal and external policy requirements
  • Enterprise-level framework to enable maturing of automation strategies
Obstacles
  • Combining the tools needed to deliver IA capability can increase tool count, cost and complexity.
  • Software engineering skills and practices are required to get maximum value from tool investments.
  • IA vendor capability expansion overlaps and confuses the tool landscape, resulting in overinvestment.
  • Steep learning curves can cause developers and administrators to revert from familiar imperative methods using scripts for infrastructure-as-code-based declarative approaches.
  • Demonstrating the return on investment (ROI) of automation projects can be challenging, especially in the short term.
User Recommendations
  • Identify existing IA tools used to catalog capabilities. Recognize use cases and document overlaps to aid decision making.
  • Assess existing internal IT skills to incorporate training needs that more fully enable IA, especially for an automation architect role to coordinate standards development and implementation. Evaluate the use of GenAI as a means to upskill teams.
  • Baseline how managed systems and tooling will be consumed (e.g., by engineers, self-service catalog, API, agentic AI or on-demand).
  • Integrate security and compliance requirements into the scope for automation and delivery activities.
  • Develop an IA tooling strategy that incorporates current needs and evolution of the near-term roadmap.
Sample Vendors
Amazon Web Services; Broadcom (VMware); env zero; Gruntwork; IBM (HashiCorp); Microsoft; Perforce Software; Progress Software; Pulumi; RackN; Scalr; Spacelift; Upbound
Gartner Recommended Reading

Product-Centric Delivery Model

Analysis By: Miriam Colman, Nabeeha Ahmed, Peter Clegg
Benefit Rating: Transformational
Market Penetration: 20% to 50% of target audience
Maturity: Early mainstream
Definition:
The product-centric delivery model features dedicated, multidisciplinary teams focused on continuous customer value through agile practices. The model is defined by its funding model, focus on customer experience, emphasis on cross-functional collaboration and distributed decision rights.
Why This Is Important
The need to respond quickly to changing market conditions and customer expectations is accelerating the adoption of the product-centric delivery model, which helps organizations align with shifting enterprise priorities. As organizations increasingly incorporate AI into products and operations, continuous learning and experimentation become more important. Product-centric teams are better suited than project teams to test changes, learn from iterations and adjust priorities over time.
Business Impact
A product-centric delivery model enables an enterprise to:
  • Focus on outcomes rather than functional outputs.
  • Improve agility in response to changing market demands and customer value prioritization.
  • Reduce silos, improve collaboration across product or customer value streams and have a flatter organization with more rapid decision making.
  • Optimize resource allocation and make operations more efficient by enabling teams to focus on specific outcomes and develop deep expertise.
Drivers
  • Organizations are increasingly focused on delivering customer value and measurable business outcomes rather than outputs or completed projects.
  • Rapid, incremental feedback is required so engineering teams can learn quickly, adjust priorities and respond to changing customer needs. Organizations need to adjust their delivery models to keep pace with market demands and increased volatility.
  • Investment and financial models must offer flexibility and facilitate evidence-based market research, aligning with and supporting corporate strategy.
  • The adoption of AI-enabled features increases the need for rapid feedback loops, ongoing experimentation and continuous tuning, which are difficult to sustain in project-centric delivery models.
  • As organizations grow and add new products or expand existing ones, they need an approach that scales effectively and allows those changes without disrupting the organizational structure.
  • Organizations need to establish clear ownership and accountability for product outcomes to ensure high-quality products and motivated teams.
  • To improve their product delivery processes, organizations need seamless collaboration across various functions.
  • Organizations require an environment that fosters innovation, encouraging the exploration of new ideas and technologies, AI in particular, to deliver cutting-edge products.
Obstacles
  • Inertia from existing organizational culture and management frameworks reluctant to disband current budgets and authority positions.
  • Difficulty overcoming change resistance and building effective product structures.
  • Lack of alignment between business and IT around outcomes, responsibilities, budgets and success metrics.
  • Lack of cohesive leadership support, which results in adoption only in pockets across the organization.
  • Outdated governance processes that incentivize control and risk aversion rather than experimentation and innovation.
User Recommendations
  • Establish clear goals and objectives for the transition and build leadership support for the necessary culture and governance change.
  • Build strong partnerships between engineering and business teams as you identify and train product managers, product owners, business leaders and team members on agile and product management practices.
  • Ensure product teams adopting AI have clear accountability for outcomes, learning and risk management.
  • Adapt governance to embrace business architecture practices such as value stream mapping, business capability modeling, and customer and employee journey mapping.
  • Move to a product funding model that allows more rapid and flexible prioritization in response to business demands and changing market conditions.
  • Track value-based outcomes of product initiatives and conduct recurring reviews to assess ongoing progress.
  • Foster a product mindset by promoting psychological safety, enabling experimentation and supporting distributed leadership.
Gartner Recommended Reading

Software Supply Chain Security

Analysis By: Aaron Lord
Benefit Rating: Transformational
Market Penetration: 20% to 50% of target audience
Maturity: Early mainstream
Definition:
Software supply chain security (SSCS) protects software, developers and environments from compromised code, tools, identities and pipelines during development, delivery and postdeployment. SSCS reduces third-party risks through policy-based dependency curation, software composition analysis (SCA) and software bill of materials (SBOM) inspection. SSCS tools establish artifact provenance and traceability via signing and verification as artifacts move through development and delivery pipelines.
Why This Is Important
SSCS transcends organizational boundaries and includes external entities in addition to internal systems. Internal systems include software delivery pipelines, software dependencies and software development environments. External entities include commercial off-the-shelf software (COTS), open-source software (OSS), third-party AI components and system images. Organizations have greater control over internal systems and little to no control over external entities.
Business Impact
  • Identifies and mitigates security and compliance risks associated with widespread third-party software use.
  • Reduces developer friction and productivity loss caused by attacks on tools, environments, pipelines and infrastructure used for software development, delivery and runtime operations.
  • Supports governance and regulatory requirements by making software delivery infrastructure auditable through automated enforcement of application security policies.
Drivers
  • State-sponsored attacks: OSS is increasingly vulnerable to infiltration by nation-state threat actors. As a result, state-sponsored software supply chain attacks have grown more sophisticated and prevalent, affecting organizations across sectors worldwide. These attacks exploit vulnerabilities in software (for example, Codecov) or compromise the software development life cycle (SDLC), as seen in the SolarWinds and NotPetya incidents. SSCS plays a critical role in reducing these risks.
  • Regulatory compliance and government mandates: Governments, policymakers and regulators worldwide mandate third-party supplier assessments, continuous vulnerability scanning and SBOMs to establish a trusted software supply chain. Examples of mandatory regulations include the Improving the Nation's Cybersecurity, EU Cyber Resilience Act, NIS2 Directive: securing network and information systems, and the Federal Food, Drug, and Cosmetic (FD&C) Act.
  • Pervasive use of open source and reliance on third-party software: Most software applications rely on third-party code through open-source dependencies. Based on hundreds of analyst interactions, Gartner estimates that more than 95% of organizations use OSS, often without full awareness.
  • Use of open-weight AI models: Easy access to open-weight large language models (LLMs) and low integration barriers introduce new software supply chain risks. The 2026 Gartner Software Engineering Survey shows that 43% of respondents rank building AI-powered features or applications among their top priorities. SSCS enables organizations to identify LLM usage, assess known model risks and enforce policies that prevent the use of unapproved models. These risks include weak model provenance, unsupported models and geopolitical restrictions on model use.
Obstacles
  • Most organizations lack a full understanding of SSCS and have not adopted a comprehensive approach to software supply chain risk. Many focus on acquiring SBOMs but have not defined how to evaluate, store, or use them.
  • Efforts to secure software artifact integrity and provenance across the supply chain are emerging but vary in scope, execution, and adoption. Policies for allowed dependencies often cause friction and require negotiation among software engineering, application security, and platform engineering teams.
  • Adoption of capabilities that harden DevOps pipelines through artifact integrity validation and automated policy enforcement remains comparatively low. Historically poor developer experiences with signing and verification workflows in continuous integration/continuous delivery (CI/CD) pipelines slow adoption. Tool heterogeneity across DevOps pipelines further complicates artifact attestation creation and pipeline integrity assurance.
User Recommendations
  • Identify and mitigate security and compliance risks associated with widespread third-party software use, including open-source and commercial software, third-party AI LLMs, Model Context Protocol (MCP) servers and containerized workloads.
  • Reduce visibility gaps in the software supply chain by using SCA and SBOMs to manage third-party risk and ensure auditability and traceability across pipeline activities and interactions in the SDLC.
  • Protect software integrity throughout the delivery process by signing and verifying build artifacts, establishing provenance data and preventing the use of noncompliant artifacts.
  • Improve the security posture of the software delivery process by automating policy enforcement across the SDLC and detecting and resolving misconfiguration errors in DevOps tooling.
Sample Vendors
ActiveState; Apiiro; Arnica; BoostSecurity; Cycode; Endor Labs; GitHub; JFrog; Lineaje; OX Security
Gartner Recommended Reading

Internal Developer Platform

Analysis By: Cary Pillers, Manjunath Bhat, Bill Blosen
Benefit Rating: Transformational
Market Penetration: More than 50% of target audience
Maturity: Mature mainstream
Definition:
An internal developer platform (IDP) is a set of tools, services, automations and information maintained as a product by a dedicated platform team. An IDP simplifies the developer experience by abstracting underlying complexity of tools and processes. Internal developer portals typically serve as the storefront for these platforms as part of a platform engineering capability.
Why This Is Important
Software engineering leaders must invest in IDPs to simplify and minimize the overhead of software engineering at scale. Digitally empowered enterprises increasingly rely on complex, heterogeneous and distributed custom-built software that rapidly changes based on customer and internal demands. However, software engineering teams often struggle to deliver these changes promptly due to poor or variable developer experience across the software development life cycle (SDLC).
Business Impact
IDPs abstract the complexity of underlying tools, services and infrastructure. IDPs also increase product teams ability to deliver customer value consistently, reliably and repeatedly. They make security, compliance and controls more consistent, and simplify the wide variety of tools used to deliver software. As part of platform engineering, IDPs improve the developer experience, thus increasing productivity while reducing employee frustration and attrition.
Drivers
  • Consistency: Automated and standardized workflows provide consistency for component creation, pipeline execution, deployment processes and AI augmented workflows.
  • Cognitive load: Adoption of modern, distributed architectural patterns and software delivery practices means that the process of getting software into production involves more tools than ever before. Agentic AI has increased the burden as developers need to know how to build, deploy and optimize AI assistants and AI agents.
  • AI governance: Enable frameworks which enable all AI generated code to be checked for compliance with organizational standards. Build in frameworks that enable AI guardrails, evaluations and observability by default.
  • Cost management: AI has increased the need for IDPs to manage and control costs. IDPs evaluate all AI-generated code for compliance with standards and enable AI guardrails, evaluations and observability to build trust in AI applications.
  • Enable security and compliance by default: Security and compliance is a high priority for all organizations. Help developers stay compliant by ensuring all code is evaluated against organizational standards. As standards evolve, update the platform to ensure developers stay compliant.
  • Need for increased speed and agility: Speed and agility are critical for CIOs when it comes to delivering customer value through software delivery. As a result, software organizations are pursuing DevOps practices through a tighter collaboration of platform and development teams to drive shorter development cycles and increase deployment frequency.
  • Scale: As more teams embrace modern software development practices and patterns, economies of scale are created, whereby there is enough value to justify creating a platform shared by multiple teams.
  • Internal developer portals: These provide the user experience for the developer to access and interact with the platform. They enable functionality like scorecards, dashboards and component management through the portal.
Obstacles
  • Internal politics: Intraorganizational fights derail platform creation. Product teams often resist giving up control of their customized toolchains. Despite its importance, improving developer experience may not be a priority.
  • Lack of skills: Platform engineering requires solid skills in software engineering, product management and modern infrastructure, all of which are in high demand but often in short supply.
  • Outdated management/governance models: Many organizations use traditional request-based provisioning models that introduce delays and complexity or do not want to give up control to enable self-service models.
  • Poor funding: Enterprises may not fund platforms without a clear ROI, which is often hard to calculate due to many unknowns.
  • Tradition of mandated platforms: Mandated platforms and DevOps toolchains with limited regard for developer experience are typically poorly adopted by developers. Thus, developers miss out on the benefits provided by the platform.
User Recommendations
  • Build strategic partnerships: Gain executive support to ensure proper funding. Engage with enterprise architecture, security, risk and compliance teams to build consensus on the right patterns and risk acceptance with the IDP.
  • Embed governance: Embed architectural guardrails, security and compliance controls into the IDP to enable placing them on the “paved road” for developers.
  • Provide a compelling experience: An internal developer portal serves as the storefront that enables self-service discovery and access to IDP capabilities. Model Context Protocol enables a new experience with integration with the platform.
  • Start small and make progress incrementally: Begin platform building using paved roads with a thinnest viable platform. Iterate on the IDP by adding new compelling experiences.
  • Treat the platform as a product: Resist the urge to treat IDP as a singular project or activity. Develop feedback loops to identify common problems and pain points to solve real problems developers face.
Sample Vendors
Cycloid; Harness; Humanitec; Mia-Platform; Northflank; Red Hat; WSO2 (OpenChoreo)
Gartner Recommended Reading

Platform Engineering

Analysis By: Neha Agarwal, Paul Delory, Cary Pillers, Bill Blosen
Benefit Rating: Transformational
Market Penetration: More than 50% of target audience
Maturity: Mature mainstream
Definition:
Platform engineering is the discipline of building and operating a self-service developer platform for software development and delivery. A platform is a layer of tools, services, automations, and information maintained as products by a dedicated platform team, designed to support software developers or other engineers by abstracting unnecessary complexity. Its goal is to optimize developer productivity, centralize governance of AI capabilities, and accelerate delivery of customer value.
Why This Is Important
Digital‑empowered enterprises increasingly rely on complex tools, processes, and AI capabilities that must evolve based on business needs. Platform engineering abstracts this complexity by providing a self‑service, curated platform aligned with developer needs and internal requirements, like security and architecture. Using platform engineering increases developer productivity, improves quality, and enables faster delivery or value.
Business Impact
Platform engineering empowers product teams to deliver software value faster. It reduces the burden of underlying infrastructure construction and maintenance, and increases teams’ capacity to dedicate time to customer value and learning. Platform engineering voluntarily standardizes the chaos associated with custom software, which results in reduced risk in security, architecture, and compliance. It helps scale and manage the complexity of AI capabilities.
Drivers
  • AI-enablement: Platforms provide paved roads, templates, and tools to enable scale AI, promote responsible AI usage, and manage the complexity of AI capabilities.
  • Cognitive load: Adopting modern, distributed architectural patterns and software delivery practices means that developing and delivering software involves more tools, subsystems, and moving parts than ever before. This approach increases the burden on product teams, which must build a delivery system in addition to the software itself. Platform engineering reduces the load of infrastructure construction and maintenance.
  • Scale: As more teams embrace modern software development practices and patterns, economies of scale emerge, which justifies creating a platform capability shared by multiple teams. This strategy is mostly of value at larger organizations where savings from platform engineering are clearer.
  • Need for increased speed and agility: The need for speed and agility of software delivery is leading software organizations to pursue DevOps, which is a tighter collaboration of infrastructure and operations (I&O) and development teams, to drive faster delivery and increased deployment frequency. This approach enables organizations to respond rapidly to market changes, handle workload failures better, and tap into new market opportunities. Platform engineering enables this cross-team collaboration.
  • Emerging platform construction tools: Many organizations have built internal platforms as homegrown efforts tailored to their unique needs, including using cloud-native application provisioning platforms and DevOps automation. Platforms generally are not transferable to other companies or sometimes even to other teams within the same company.
  • Emerging developer portals: A healthy internal developer portal market is enabling front-end platforms, and Backstage is a popular open-source solution.
  • Infrastructure modernization: Modernization efforts push I&O and development teams to adopt platform engineering to deliver more value to the business.
Obstacles
  • Platform engineering is easily misunderstood: Traditional models of mandated platforms and DevOps toolchains can be relabeled and not achieve the true benefits of platform engineering.
  • Lack of skills: Requires high in-demand skills of software engineering, product management, and modern infrastructure.
  • Outdated management/governance models: Reliance on request-based provisioning models can create delays and complexity.
  • Internal politics: Intraorganizational fights, no appetite to improve developer experience, or resistance from teams to give up their customized toolchains stalls efforts.
  • Funding: Enterprises may refuse to fund platform engineering without a clear ROI. Attribution of the costs to user budgets is also tricky. Measuring the benefits/outcomes is critical.
  • Scaling: It’s hard to scale and meet the constant need for platform evolution.
  • Innovation: Platform engineering has not kept up with the pace of AI and internal developer portal innovation has stagnated.
User Recommendations
  • Begin by setting up a platform team with aligned stakeholders and a product owner to guide platform-building efforts with the thinnest viable platforms for the complex infrastructure underneath cloud-native and distributed applications, including technologies like containers and Kubernetes.
  • Enable shift-left and shift-right security within DevOps pipeline platforms, which will provide a compelling, paved road to engineers.
  • Embed architectural guardrails, compliance controls, and any other nonfunctional requirements into the platform to further pave the road for developers.
  • Refrain from expecting to buy or build a complete platform, as it is unlikely that any commercially available tool will provide the entirety of the platform you need.
  • Implement an internal developer portal, which enables self-service discovery and access to internal developer platform capabilities. Consider the Backstage open-source solution, if resources permit, or other commercial tools.
Gartner Recommended Reading

Appendixes


See the previous Hype Cycle: Hype Cycle for Platform Engineering, 2025

Hype Cycle Phases, Benefit Ratings and Maturity Levels

Hype Cycle Phases

Phase
Definition
Innovation Trigger
A breakthrough, public demonstration, product launch or other event generates significant media and industry interest.
Peak of Inflated Expectations
During this phase of overenthusiasm and unrealistic projections, a flurry of well-publicized activity by technology leaders results in some successes, but more failures, as the innovation is pushed to its limits. The only enterprises making money are conference organizers and content publishers.
Trough of Disillusionment
Because the innovation does not live up to its overinflated expectations, it rapidly becomes unfashionable. Media interest wanes, except for a few cautionary tales.
Slope of Enlightenment
Focused experimentation and solid hard work by an increasingly diverse range of organizations lead to a true understanding of the innovation’s applicability, risks and benefits. Commercial off-the-shelf methodologies and tools ease the development process.
Plateau of Productivity
The real-world benefits of the innovation are demonstrated and accepted. Tools and methodologies are increasingly stable as they enter their second and third generations. Growing numbers of organizations feel comfortable with the reduced level of risk; the rapid growth phase of adoption begins. Approximately 20% of the technology’s target audience has adopted or is adopting the technology as it enters this phase.
Years to Mainstream Adoption
The time required for the innovation to reach the Plateau of Productivity.
Source: Gartner

Benefit Ratings

Benefit Rating
Definition
Transformational
Enables new ways of doing business across industries that will result in major shifts in industry dynamics
High
Enables new ways of performing horizontal or vertical processes that will result in significantly increased revenue or cost savings for an enterprise
Moderate
Provides incremental improvements to established processes that will result in increased revenue or cost savings for an enterprise
Low
Slightly improves processes (for example, improved user experience) that will be difficult to translate into increased revenue or cost savings
Source: Gartner

Maturity Levels

Maturity Levels
Status
Products/Vendors
Embryonic
In labs
None
Emerging
Commercialization by vendors
Pilots and deployments by industry leaders
First generation
High price
Much customization
Adolescent
Maturing technology capabilities and process understanding
Uptake beyond early adopters
Second generation
Less customization
Early mainstream
Proven technology
Vendors, technology and adoption rapidly evolving
Third generation
More out-of-box methodologies
Mature mainstream
Robust technology
Not much evolution in vendors or technology
Several dominant vendors
Legacy
Not appropriate for new developments
Cost of migration constrains replacement
Maintenance revenue focus
Obsolete
Rarely used
Used/resale market only
Source: Gartner

Evidence


1 Gartner Software Engineering Content Survey for 2026. This survey was conducted to provide a comprehensive understanding of the current landscape in software engineering, as well as to determine the priorities and strategic challenges of software engineering leaders. It also aims to identify the demand for various roles and skills within software engineering organizations, and assess their budget expectations, team structures and organizational outcomes. Finally, it explores the integration of AI in software engineering workflows and its impact on engineering organizations. The survey was conducted online from August through November 2025 among 482 respondents from the U.S. (n = 360) and U.K. (n = 122). Qualifying organizations operated in multiple industries and reported enterprisewide revenue for fiscal year 2024 of at least $250 million or equivalent. Qualified participants were highly involved in managing software engineering/application development teams and the activities they perform. Disclaimer: The results of this survey do not represent global findings or the market as a whole, but reflect the sentiments of the respondents and companies surveyed.