facebook

The “Yesterday” Problem: What I’ve learned from customers who challenged the traditional procurement timeline  

By Jag Gattu, CEO, UptimeAI  

This blog was originally published on LinkedIn.

 

I have this question I ask almost every operations leader I meet: when do you actually need this running? I ask it half out of curiosity and half because I already know the answer. Nine times out of ten, it’s some version of “yesterday.” Somebody’s losing sleep over a unit that keeps tripping, or a maintenance backlog that never shrinks, or a retirement that’s about to walk two decades of institutional memory out the door. There’s real urgency in the room.  

And then, almost every time, that urgency hits a wall when it comes to the procurement process. First, they’re waiting on legal. Next, it’s IT. Someone has to find out who actually owns all the systems that need to be connected together. And by the time all of that has happened in sequence, the six months that everyone swore they didn’t have at the start of the project have quietly gone by.  

I want to be careful here, because I don’t think the lesson is “move faster at all costs.” I’ve watched companies get burned by exactly that instinct — skipping a scoping session or waving off a data question can add more delays later. Whatever we build, it has to earn its way into a plant’s operations, and that takes real diligence. What I’ve learned instead, from watching a lot of these cycles up close, is that diligence doesn’t have to walk a straight line. Many procurement tasks can occur simultaneously; they just usually don’t because it’s not the way things have historically been done.   

Here are a few of my favorite techniques that our customers have used to accelerate the value generation of their AI project by choosing not to wait on traditional timelines.   

1. Be an advocate for imperfect data  

I remember a call with an operations director at a cement plant who was clearly ready to move, but his team kept getting stuck arguing over whether they had “enough” and “good enough” data. Every week someone identified a new handful or tags, or a new sensor that would be “nice to have when trying to predict XYZ.” It could have gone on indefinitely.  

At some point he just made a call: they’d agree on the handful of failure modes and the data around the assets that actually mattered for proving the case, then treat everything else as things that could be done in parallel once the solution was live. That one decision point—to accept that some things would be ready now, others would improve over time, and even more opportunities would be uncovered by their use of the technology—avoided weeks of delays in getting the software operational.  

2. Know your scope, your systems, and your stakeholders  

We worked with a project manager for deployment at a large refining client whose team showed up on day one with a clearly defined pilot scope, the KPIs they’d measure success on, the data they’d need to deliver successful results, and the people with the right access to the relevant systems to make it all happen.   

This might sound trivial, but I saw another engagement get delayed by over a month because a new data source was added to the scope which lived at a different level of the enterprise architecture and had a different owner than our other refinery data sources. That owner wasn’t up to speed on our project, and we were fighting for their time with other priority projects that they’d known about for much longer.   

3. Identify, and challenge, dependencies 

One of our early power generation customers taught me something I hadn’t fully appreciated: a lot of the delay in these deals comes from waiting on others out of habit rather than requirement. 

Their security and architecture review ran at the same time as the final commercial conversation, not after it, because somebody on their team asked “why would these two things need to happen in order?” and nobody had a good answer. Procurement and legal moved on a similar clock. Nothing about the actual approvals changed — same reviewers, same rigor. What changed was that they finished around the same time the contract did. That business was ready to execute the week they signed, instead of six or eight weeks later, which is close to how long that kind of review usually takes when everyone assumes it must come last.  

What I want you to take away from these stories 

None of these customers moved faster by cutting anything out. They moved faster because they stopped assuming that due diligence must happen in series. Confirming the use case, mapping the data, starting the security and legal review, building the onboarding plan — almost none of it actually depends on each other.   

If your team is telling you they needed this yesterday, I’d take that seriously — not as a reason to skip anything, but as a reason to ask, honestly, which parts of the next few months are actually sequential out of requirement, and which ones are just habit and can be accelerated to help get them the solution they need, yesterday.  

I’m grateful to the customers who have proved to me that it’s possible to go from demo to go-live in less than a quarter. We’re taking those learnings, applying them to every customer sale, and helping our new customers push to the frontier of adoption of AI reasoning agents for industrial operations.   

 

Our Commitment to Customer Value: Why We Asked Someone Else to Grade Our ROI

By Jag Gattu, CEO at UptimeAI 

 

This article originally appeared on LinkedIn.

 

Every industrial software vendor promises transformation: faster diagnoses, fewer surprises, a healthier bottom line.  

Today’s market places too much responsibility on the customer when it comes to validating these promises. For executives deciding where to put scarce digital investment dollars, the ambiguity gets expensive, and it can be the reason that a promising pilot never reaches enterprise scale. It’s not because technology isn’t valuable, but you need to be able to build a business case that finance can trust. 

This is why, in a market swamped with self-reported ‘proven results’ we invested in helping your teams build a confident business case. Conducting economic studies like the Verdantix Verified Value Delivery (VVD) methodology is not inexpensive and it’s not trivial. It involves hours of interviews with UptimeAI customers to parse out exactly how they are realizing value from the technology, then building that into a calculation framework that can apply across company sizes, industry verticals, and products deployed.  

We think it’s important that our customers understand what value they should expect from the earliest days of engagement – in the form of proof, not promise.  

How do you measure the impact of something that never happened? 

One of the biggest challenges for predictive technologies, particularly in the reliability space, is the definition of success. When our software works exactly as intended, the outcome is often nothing happens. The cost savings of avoiding unplanned failures and unit shutdowns are typically only quantified when the shutdown actually occurs. So how do you quantify the value of something that never happened? 

Since UptimeAI started, our customer success team has worked with every customer to track every alert and diagnosis that prevented a failure and aligned with the customer on what that catch was worth. They comb through past failure events, the lost production, the maintenance expenses, then the system learns the value in warning of that failure mode, so that next time value is automatically assigned. We know that our champions are constantly being asked to justify spend on operational technologies, and we view it as our job to make that task as easy as possible for them.  

Customer interviews confirmed UptimeAI’s commitment to defensible value 

Verdantix conducted interviews of 5 global customers, using their responses to build a model of expected returns for various company demographics. Specifically, for a model $200M-revenue manufacturing site, 3-year ROI of 197% grows to 250% as they expand enterprise-wide, and the software has fully paid for itself in 11 months. In some cases, it’s much faster.  

For one major global cement company – “We had 4 critical alerts in first 6 months of the pilot, when we were very conscious of the value to justify further investment in the solution, and this recouped the cost of the investment.” We continue to compress this payback period by delivering more value, sooner, through more agent products and workflows. 

An exciting future ahead 

197% ROI is exciting, but that’s just getting started. In just 6 months since the study was completed, there are already so many additional value streams that our customers are seeing leveraging new agents. I can’t wait to see what the next year will bring for our current and future customers. 

 

For the full details, read the full Verdantix VVD study.

Build vs. Buy, Part II: The Very Real Barriers to BIY in Highly Regulated Industries

By Jag Gattu, CEO, UptimeAI 

This post originally appeared on LinkedIn.

 

My first post on this topic highlighted the shift we’re seeing in how many tools we build in house vs. buy, and where we draw the line. At UptimeAI, our line sits at the transition point from retrieval to reasoning.  

I took a step back to see where others in the market are drawing the line. While I found general excitement about the potential of BIY (Claude Code and the likes), I found that leaders are being appropriately calculated in their consideration of these projects. There isn’t just one line that they’re drawing for when to build versus buy a solution, but many, and the lines link back to the same challenges that have plagued digital projects in our industry for the past decade.  

Here are the three biggest challenges with BIY agentic AI projects in the process industries.  

1. Governance

In a recent blog post, Verdantix highlighted that AI sprawl is becoming a C-suite problem. Sprawl is a second-order cost to unlimited BIY development. Every vibe-coded agent is also one more tool requiring governance — and a new source of shadow AI, data leakage, and regulatory exposure if it’s not accounted for. Even Amazon has acknowledged internally that “the AI boom is driving duplication and fragmentation, not less of it.” Building a tool and then governing it responsibly, indefinitely, in a live plant are two different jobs, and the governance doesn’t get easier just because the build got faster. Industrial software companies have been operating in this area for decades and have the data governance and cybersecurity infrastructure already in place to deploy agentic AI projects fast without risk. 

2. Scale

According to McKinsey’s 2025 State of AI report, less than 10% of organizations currently experimenting with building AI agents are scaling them. BIY pilot development is easier than production deployment. You can have a demo working reliably on clean, well-behaved offline data in a matter of hours. Standing up something that holds up to the challenges of real operations data — fragmented, inconsistently modeled ERP, MES, historian, and SCADA data, spanning reliability, process, and operations teams — is another. Agents built on top of fragmented data don’t fix fragmented decision-making, they automate and amplify it. That’s true whether you built the agent or bought it. The difference is whether you’re building the curation, integration, and verification layer from scratch, or standing on one that’s already been proven across other plants like yours. 

3. Impact

In a recent LNS Research blog post, Research Analyst Vivek Murugesan emphasizes the importance of “leading with business problems and not technology itself.” Are you looking for a personal productivity tool or a way to make margin-impacting operations decisions? A tool that surfaces the right document or summarizes a work order can be easily built, and it is useful, but it isn’t the same as a system that provides operators with confident decisions to high-stakes calls. Margin is lifted through improvements to availability, quality, and reliability. Personal productivity can support improvements to these areas, but it won’t move the needle on margin the way the components of operational excellence can. An enterprise deployment of a Root Cause Agent, for example, requires automated expert reasoning, robust data handling, and continuous learning and improvement loop. That’s not retrieval, it’s judgment, and it’s proven to provide rapid technology ROI through improvements to diagnosis time, accuracy, and avoided downtime. 

Draw YOUR right lines 

The challenges raised above aren’t reasons to avoid BIY altogether. There are problems where in-house solutions are sufficient, like our sales enablement tool I mentioned in my last post. Retrieval has become easy to build and scale in-house. But when making build v. buy decisions around reasoning and judgement—the agentic AI capabilities required to make impactful decisions in a regulated industrial environment—trusted partners can effectively overcome the governance, scale, and impact hurdles. 

The Stories Behind the ROI: What 197% Looks Like on the Plant Floor

By Shruti Kela, Engagement Manager at UptimeAI 

 

This blog post originated on LinkedIn.

 

I have spent my career on one question from both ends: what is the real business impact of the work that keeps a plant running. At SLB, I worked in the field, close enough to the equipment to know what a failure costs when it lands. At McKinsey, I built the business cases that decided whether a plant spent millions to prevent that or lived with the risk. Knowing the price of a breakdown is what makes a prevented one so valuable, and also what makes it so hard to prove. The biggest wins in this business are the failures that never happen, and you cannot show a board a breakdown that was quietly avoided.

Unless you catch it in the act. UptimeAI flags a failure early, names the likely cause, and recommends the fix, so engineers act before the loss lands, and every catch is logged as it happens. The breakdown stays invisible, but the savings do not, and those records are exactly what Verdantix set out to audit when UptimeAI brought them in to conduct a Verified Value Delivery study.

That kind of proof is rare. In their 2026 Global Corporate Survey, Verdantix found that 85% of industrial firms say measuring the ROI of their AI tools is a real barrier, which is how promising pilots quietly die. Working from the records of five UptimeAI customers, their independent study landed on a 197% three-year return for a typical site, growing toward 250% at scale.

A high-level metric like that is only as trustworthy as the data that sits under it, so instead of the top-down math, I sat down with my colleagues on UptimeAI’s Customer Success team to sum up the individual moments that compound into triple-digit ROI.

The building blocks for triple-digit ROI

A failure with no alarm

At a cement plant, the kiln is the whole business. Every ton of clinker produced runs through it. One of the sneakiest ways to lose output is coating building at the kiln inlet, and there is no alarm for it, because it never shows up as one bad number. It shows up as small shifts across many at once, pressure edging up, gas composition drifting, each too minor to notice alone. Operators cannot watch every point at once. UptimeAI read them as a system, recognized the coating pattern, and flagged it hours ahead with a clear instruction: ease off the fuel, check the chemistry, prepare to clean. The team cleared it during a short, planned stop instead of a reactive multi-day shutdown, and kept the kiln producing. The value of this catch was amplified a few weeks later when the same issue occurred on a kiln at another plant; six figures and dozens of production hours were kept rather than lost.

One flat gauge hiding many moving ones

A gas turbine at a power operation was holding steady at full load, nothing an operator would look at twice. Underneath, UptimeAI saw what no single gauge would: temperatures rising together across the bearings, the wheel space, and the exhaust, all while output stayed flat. Read one by one, nothing alarmed. Read together, they pointed to one cause. UptimeAI traced it to an inlet guide vane that had drifted open, pulling excess air through the compressor and loading the bearings, and pointed the team straight at the vane hardware. They inspected it, found heavy wear, and planned the fix on their own terms instead of losing the machine to a trip. The turbine kept generating the whole way through. Value the customer confirmed: $280,000.

Diagnosing the cause, not the symptom

When a pump at an oil and gas plant started shaking, everyone looked at the bearing. That is where the vibration showed up, so that is where a normal investigation goes. UptimeAI’s Root Cause Agent looked wider. It read dozens of signals together, pulled in the plant’s own maintenance records and years of documents, and built a causal chain in minutes. Its top answer, at 90% confidence, had nothing to do with the bearing. A seal job weeks earlier had been aligned cold, and once the pump warmed up it pulled out of true. The agent even surfaced a write-up on a sister machine with the same story. The team corrected the alignment instead of tearing into the bearing, kept the unit running, and avoided a failure worth more than $500,000. The fix was never where everyone was looking.

When the smartest answer is “it’s not broken”

Not every alert is a real problem, and the safe reaction, shutting down to go look, is exactly how you lose production you never needed to lose. At an oil and gas operator, a bearing alert fired, and the obvious move was to plan a repair and take the unit down. UptimeAI’s Root Cause Agent challenged it, worked through the data, and found the bearing was fine; the instrument reading it was faulty. It cleared several instrument faults in minutes. The team kept the unit running and skipped an outage and a repair it did not need, worth about $750,000.

Right-sizing maintenance, not just cutting it

At a large power plant, the preventative maintenance plan was set by equipment class and almost never revisited. Every pump on the same oil-change clock, every bearing swapped on the same calendar. Revisiting it meant weeks of expert time pulling work histories, so it rarely happened. UptimeAI’s Maintenance Optimization Agent analyzed the plan continuously and provided ranked recommendations to minimize spend across preventative and corrective maintenance. On a single pump, it found four moves at once: do more where failures were slipping through, less where the data proved it safe, add a missing task, and drop a calendar task that live sensors already covered. UptimeAI pushed approved changes straight into the CMMS, no re-keying. One pump surfaced tens of thousands in savings. Extended across the site, close to $300,000, without a reliability engineer touching a spreadsheet.

A capability that sits inside the customer’s team

The real test of a monitoring tool is whether the customer’s own people run it. At a major North American chemicals producer, the engineers approved and closed most alerts themselves, requiring minimal support from the UptimeAI team. When they wanted a second opinion, they asked UptimeAI’s GenAI copilot, Rooty. A pair of catches on a critical pump went onto the maintenance plan before either failed, so the unit kept running. When the team wanted to test UptimeAI’s predictions against their own historical data, it flagged real past failures months ahead of when they actually happened. With UptimeAI assisting them, the site team caught and fixed problems before they ever became problems.

The math underneath

Every one of these stories represents a single line item in the Verdantix calculation of 197% return. The full study lays out where each dollar comes from, an 11-month payback, and how the return climbs toward 250% as the technology is rolled out across plants or sites. If you have ever watched a promising pilot stall because no one could prove what it was worth, a read of the full study is worth your time.

Get the full Verdantix Verified Value Delivery study

Build vs. Buy: Advice for Operations and Digital Executives in the Age of Generative AI

by Jag Gattu, CEO at UptimeAI

 

This article originally appeared on LinkedIn.

 

If you’re a CIO, CDO, of VP of Operations and you aren’t asking which of your SaaS tools could be retired in the age of generative AI, you’re missing a big part of your job right now. 

Here’s a small example from our own experience. Last week, one of our team members used Claude Code to rebuild a sales relationship-mapping tool we’d been paying for as a SaaS subscription. A few days of work, and we no longer need to renew that contract. Genuinely great use of GenAI — a well-scoped, single-purpose application, recreated in-house faster and cheaper than re-negotiating the contract with the vendor. 

Here’s what we didn’t try to rebuild: our CRM. Nobody on our team seriously proposed it, and not because it wasn’t technically possible. It’s because scope, scale, complexity, and the sheer depth of integrations a CRM touches make it a terrible build candidate — even in a world where a single engineer with the right AI tools can do more than an entire team could five years ago. Add in how fast the economics of the big model providers are shifting, and “build it ourselves” gets shakier by the month, not sturdier. 

That contrast is the whole build-versus-buy question in miniature. GenAI has genuinely moved the line on what’s worth building in-house, but it hasn’t erased the line. 

And the data backs this up. MIT’s widely-cited 2025 study (enter “95% of AI pilots fail to make it to production”) on enterprise AI found that internally built AI deployments succeed at roughly half the rate of externally partnered ones — 33% versus 67%. Gartner has separately warned that at least half of generative AI projects will blow through budget due to poor architectural choices, and that most organizations attempting to build custom models will eventually abandon those efforts due to cost, complexity, and technical debt. That’s not a knock on any one company’s engineering team. It’s what happens structurally when you take on a mission-critical or business-critical system that has to keep working, keep learning, and keep being right, indefinitely. 

So where does something like UptimeAI fall on that line? For the majority of industrial organizations, we see a strong case for not building it yourself — but it’s worth being precise about why, because it’s no one reason between scale, complexity, risk, and upside.  

Most organizations that try to BIY this space end up building a knowledge graph, wiring up retrieval over their operational data, and putting a chatbot on top. That’s a real project, and a legitimate one — but it’s retrieval. It answers, “what does the data say.” 

Bridging from retrieval to reasoning is a big leap. Ours come with the data access, the knowledge graph, the contextual relationships between assets and failure modes, the domain skills of experienced engineers, and the orchestration across sub-agents already built in — tuned to respond to specific high-impact business problems like “what’s actually wrong, and what should you do about it” the way your best expert would. That’s not a chatbot with good retrieval. That’s domain expertise and engineering judgment, encoded, and scaled. It’s a much harder thing to stand up yourself, and it’s exactly the layer where buying beats building. 

Retire the SaaS tools GenAI can genuinely replace. Just don’t confuse that win with being able to BIY it all. Reasoning and retrieval are fundamentally different jobs — and at UptimeAI we take pride in doing the hard jobs in a way that generates big returns for your business. 

 

Schedule a demo to see our off-the-shelf reasoning agents for yourself. 

HAZOP Analysis Agent Reduced SME Time Burden by 60% for North American Chemical Manufacturer

The Challenge: HAZOPs Drained Valuable SME Capacity

As required by OSHA, a mid-sized North American chemical producer conducted Hazard & Operability (HAZOP) studies for each unit in the plant on a 5-year frequency. But with 10 process areas and a continually shrinking pool of SMEs to pull in, HAZOP workshops were a major draw on the site’s most experienced personnel.

Every revalidation cycle required process engineers, operators, EHS leaders, and maintenance experts to spend weeks walking through every deviation-guideword row, often searching for missing information or rebuilding process context rather than evaluating risk. The site was seeing a downward trend in key business metrics like product quality, downtime, and maintenance cost. The people they counted on to improve these metrics were trapped in a conference room highlighting P&IDs.

The Solution: Auto-generated HAZOP Drafts Built on Plant Documentation & Expertise

To reduce the preparation burden on SMEs, the company deployed UptimeAI’s HAZOP Analysis Agent to automate the most time-intensive portions of the PHA process. Starting from existing P&IDs, the agent automatically identified equipment, instrumentation, and process flows to build a digital representation of each unit. The agent generated draft HAZOP nodes, organized relevant documentation, and produced initial deviations, causes, consequences, safeguards, and recommendations aligned to the company’s risk framework.

Rather than starting from a blank worksheet, SMEs entered workshops with a fully populated draft analysis already in place. This shifted their role from manually assembling information to validating risk scenarios, refining recommendations, and making higher-value operational decisions.

The Impact: More Time Spent Managing Risk, Less Time Preparing for Reviews

By deploying the HAZOP Analysis Agent, the site significantly reduced the operational burden associated with recurring PHAs.
Hazop Agent - UptimeAI

Results That Scale

After successfully demonstrating the impact of HAZOP Analysis Agent on SME time for one process unit, the site rapidly expanded to include all site process units + sitewide deployments for all other plants under this business unit. The sites are leveraging the unlocked SME capacity for improvement projects targeting process efficiency, reliability, and quality metrics, while the process safety incident rate is at the lowest it has been in the last decade.

UptimeAI is where the market is heading

By Jagadish Gattu, CEO of UptimeAI 

This article first appeared on LinkedIn.

Verdantix biennial Green Quadrant for Asset Performance Management (APM) was recently released, and this version was different. AI was no longer treated as a singular capability but woven into the thread of every evaluation criteria. As the newest company in a market dominated by incumbents we saw this as an advantage. Being born in the age of AI means our AI products were built that way from the ground up, not sprinkled on top of decades old technology. This advantage was reflected in various capability scores, and also the overarching narrative of the report.  

The No. 1 Score in Market Vision & Business Strategy 

The momentum (x) axis in a Green Quadrant correlates strongly to the size of the business and dominance in the marketplace, which can be a big advantage for legacy companies, who have had decades to grow sales and following to where they are today. For newer companies, the place to shine amidst the momentum criteria is less about where you’ve been and more about where you’re going.  

When I saw that UptimeAI had been awarded the highest score in the field for Market Vision & Business Strategy (a 2.9 on a 3.0 scale), I was not surprised. In my article on closing the gap between insight and execution, I discussed the #1 takeaway from the 2025 Verdantix Asset Management Council. The 13 asset management leaders from major energy and industrial companies came together and deduced that: 

Industrial agility – the ability to rapidly adapt operations, processes and workforce focus – is becoming increasingly relevant because of macroeconomic pressure and increasing dislocation. Improving agility by closing the gap between insight and execution will be the key to enhancing operational excellence amidst reskilling, data-fragmentation and scaling challenges. – Verdantix 2025 Asset Management Council

This was a recurring theme in the 2026 Green Quadrant, and a common strength amongst the companies that scored highest in vision and strategy.  

Agentic AI is transforming APM from insights-driven to action-oriented 

Past definitions of APM held up predictive analytics capabilities as the gold standard for uncovering insights hidden in untapped data. But today, the data’s been tapped, the insights are piling up, yet the outcomes still lag. It turns out it was never a shortage of insights limiting our industry’s margins. It is the shortage of expert decisions that drive actions and the subsequent outcomes that we’ve been limited by.

In talks at CERAWeek and other events, I’ve described the expert decision  bottleneck that’s costing industrial organizations millions every year. Predictive analytics stops at the point of detection, awaiting expert interpretation to arrive at a decision. The amount of time spent between uncovering an insight and getting to an optimal decision is the decision latency created by the expert bottleneck. Decision latency costs organizations millions in failures, repairs, unplanned downtime, and excessive preventative maintenance. Overcoming that expert bottleneck is the key to creating an APM program that delivers real margin impact. 

Decision latency in industrial operations
The expert bottleneck created by decades of tools that have focused on insights.

UptimeAI takes an agent-first approach, applying AI to emulate key maintenance workflows such as optimization and root-cause analysis. It continuously evaluates operating conditions, asset criticality, failure mode and effects analysis (FMEA), work orders and equipment documentation to identify underlying failure drivers, optimize maintenance strategies and recommend actions, with transparent reasoning behind each decision. This shift towards agentic AI also raises the bar for incumbents: success is increasingly dependent on either innovating quickly or forming partnerships to effectively leverage agents. – 2026 Verdantix Green Quadrant: APM

When the decision gap closes, results compound 

One of the strengths of the Verdantix research process is their use of customer interviews to validate market presence and product capabilities. Speaking with UptimeAI customers, Verdantix verified maintenance time savings of 70 hours per month and annual reliability + performance savings of $10M per year at a single coal plant for one of India’s top 5 largest power generation companies. In a second interview with a top 10 US cement manufacturer, Verdantix confirmed $500k savings and 24h of avoided kiln downtime from the early diagnosis and mitigation of a single anomaly event.  

When the gap between insights and actionable decisions collapses, margin growth accelerates. UptimeAI has delivered repeatable results across global industry leading organizations in oil and gas, cement, chemicals, and power generation.  

The most forward-looking industrial organizations are already making the shift — and UptimeAI is making it possible. 

Expanding Expert Decision Capacity in Energy and Heavy Industry with AI Reasoning Agents

By Cody Berra, Senior Solution Consultant at UptimeAI 

This piece was originally published on Smart Industry.

Over the past decade, energy and heavy industrial companies have invested heavily in digital transformation. Process historians capture millions of data points per asset, maintenance systems track decades of work history, and engineering knowledge is stored across documents, drawings, and reports. Despite this, productivity gains have been modest.  

The issue is no longer visibility, it is decision throughput. Across operations, maintenance, and reliability, there is no shortage of data or even insights. What is limited is the ability to consistently interpret that information and translate it into the right action at the right time. 

Decision latency in industrial operations
The decision bottleneck created by a shortage of industrial experts is limiting decision velocity, directly impacting margins.

Every day, engineers are asked to answer questions that require connecting multiple domains. Is this vibration issue mechanical or driven by upstream process conditions? Should equipment be taken down now, or can it run until the next outage? Are current conditions within a safe operating envelope, or are small deviations stacking up into a larger risk? These are not simple questions. They require context, experience, and judgment. 

In most plants, that capability sits with a small number of experienced engineers. They pull data from multiple systems and piece together conclusions manually. This work is time intensive and difficult to scale, which means many decisions are delayed or never made at all. As assets age and experienced workers retire, that constraint becomes more pronounced. The problem is less about data availability and more about scaling expertise. 

From Detection to Decision 

AI reasoning agents are emerging to address this gap. Unlike earlier systems that focused primarily on detecting anomalies, these technologies are designed to replicate how experienced engineers diagnose problems and make decisions. They bring together time series data, maintenance history, and engineering context, then apply domain specific reasoning to connect symptoms to likely causes and recommended actions. 

Instead of simply flagging a deviation, the system produces a structured explanation that outlines what is happening, why it is happening, how confident the conclusion is, and what actions should be considered. This shift from detection to decision support allows organizations to act more consistently and with greater confidence. 

Use Case 1: Root Cause Analysis on Rotating Equipment 

A common example can be found in rotating equipment. Consider a centrifugal pump that begins showing elevated vibration following a maintenance event. A traditional system will flag the anomaly, after which an engineer investigates: reviewing trends, checking maintenance history, and consulting documentation. Depending on complexity, this process can take hours, or even days. 

A reasoning agent specializing in root cause diagnosis and correction compresses that workflow. It can automatically correlate the vibration increase with a recent coupling disassembly, evaluate patterns consistent with different failure modes, and surface similar historical cases on comparable equipment. When it classifies this event as misalignment, rather than other failure modes, it draws from past data and experience to prescribe a laser alignment with hot thermal growth targets, a soft-foot check, and a revision to PM procedures. 

While the agent provided earlier detection than legacy systems, the bulk of the value came from having a faster and more consistent diagnosis. Plants using this approach are reducing time to resolution and avoiding repeat failures by addressing underlying causes rather than reacting to symptoms. 

Use Case 2: Maintenance Optimization in Practice 

While root cause analysis addresses individual events, maintenance strategy presents a broader challenge. Many organizations still rely on time based preventive maintenance, where equipment is serviced at fixed intervals regardless of condition. Over time, this leads to unnecessary work on healthy assets and missed failures on assets that degrade between intervals. 

A maintenance optimization agent introduces a continuous feedback loop. It analyzes historical work orders, failure events, and operating conditions to determine how maintenance frequency impacts reliability for each asset. Rather than applying a uniform strategy across an asset class, it evaluates equipment based on its actual operating history.  

Heavy industry using AI reasoning agents to optimize industrial operations
Image. A reasoning agent does continual PM interval optimization, weighing PM v. CM costs and recommending changes with cost basis.

For example, a plant may perform quarterly maintenance on pumps yet continue to experience recurring failures. The system can quantify the relationship between maintenance intervals and failure rates, helping determine whether the issue is insufficient maintenance or, in some cases, excessive maintenance that introduces risk. Each recommendation is supported by a clear cost and risk trade off, outlining expected changes in failure frequency, maintenance cost, and potential production impact. 

Engineers can test scenarios, apply constraints, and review assumptions before implementing changes. Over time, this shifts maintenance strategy from a static, experience-driven practice to a dynamic, evidence-based process, allowing teams to continually re-focus efforts towards the areas of greatest impact. 

Use Case 3: HAZOP as an Ongoing Capability 

Process hazard analysis (PHA) is critical, regulatory activity to maximize process safety performance, but it’s traditionally been static. The most common format of PHA, Hazard & Operability (HAZOP) studies are conducted on a 5-year cycle, with results captured in documents that are difficult to access and rarely used in daily operations. 

A reasoning agent for HAZOP efficiency changes both the speed and frequency of this work. By ingesting P&IDs and engineering documents, the system builds a connected model of the process and generates a structured HAZOP draft, including nodes, deviations, causes, consequences, and safeguards. What once required weeks or months of preparation can now be generated in days, allowing engineers to focus on analysis rather than assembling information. 

More importantly, the analysis becomes more consistent. Instead of relying on what a team can recall in a workshop, agents can systematically evaluate deviation scenarios across the full process, including interactions that span multiple units. Engineers still review and refine the output, but they begin from a well-structured and evidence-based starting point. 

The result is a more consistent, more accurate approach to process safety. Rather than waiting for the next revalidation cycle, teams can revisit the full HAZOP analyzes when operating conditions change or equipment is modified, minimizing gaps between design assumptions and actual operation. 

Measurable Impact and the Path Forward 

Early adopters across oil and gas, chemicals, and power generation are seeing a positive impact on operations with the implementation of reasoning agents. These include earlier detection and diagnoses that enable planned mitigations at the source of process or equipment issues, reduced maintenance costs through better targeting of work, and improved asset performance. In many cases, teams are also seeing gains in energy efficiency, particularly in industries where small deviations carry significant cost impact. 

Equally important is how expertise is deployed. Experienced engineers are no longer consumed by routine troubleshooting. Instead, their knowledge is applied more broadly across the organization, supported by systems that make their reasoning repeatable. 

AI reasoning agents are not a replacement for human expertise, they are an extension of it. By making diagnostic workflows and decision logic scalable, they allow organizations to apply expert level thinking across more assets for more decisions. For an industry facing aging infrastructure, workforce constraints, and increasing pressure on margins, these agents chart a path towards operational stability and business longevity. 

The Gap Between Insight and Execution: Where Industrial Performance Is Won or Lost

By Jagadish Gattu, CEO of UptimeAI 

This article originally appeared on LinkedIn

The number 1 takeaway from Verdantix 2025 Industrial Asset Management Council reflects something we hear from our customers every week:

Industrial agility—the ability to rapidly adapt operations, processes and workforce focus—is becoming increasingly important, and the best way to improve industrial agility is to close the gap between insight and execution.  

Industrial agility isn’t hindered by visibility; it’s a velocity issue. More precisely, it’s limited by the ability to turn an insight into a decision, and a decision into an action before the optimal moment passes.

Decades old challenges with a new urgency 

The council members discussion highlighted tensions that will sound familiar to anyone running asset-intensive operations: balancing quality standards against delivery commitments, managing aging assets against cost pressure, maintaining continuous runtime in markets that keep getting less predictable. These aren’t new challenges, but with a slew of external factors complicating margin equations, there’s a new urgency.  

Historically, the workflows around industrial agility have been software initiated but human bottlenecked. Industrial companies have the data necessary to solve most problems. Many of our customers even have existing software to detect when operating conditions have shifted. But then those detected events await human expert interpretation before a decision is made, and an action initiated.  

Decision latency is introduced when software detects an issue, but then it takes days or weeks of human response time to act on the insight. Decision latency is the killer of industrial agility.  

This decision latency created when you have an overabundance of software that detects and an underabundance of human experts that decide is why we built UptimeAI, and why this finding hits close to home. 

Early industrial AI tools created an insight-to-action gap

The first wave of industrial AI solutions focused largely on prediction.  

  • Can we detect an anomaly earlier?  
  • Can we forecast a failure before it occurs?  

Those capabilities are genuinely valuable, and they’ve become increasingly valuable as they’ve matured over the past several years. But prediction alone doesn’t close the insight-to-action gap. A flagged anomaly that doesn’t result in an automated, accurate diagnosis, and mitigation hasn’t actually moved the needle. It just adds another alert to a queue that a human still needs to reason through manually. 

The next generation of industrial AI presents organizations with the opportunity to no longer stop at the alert. AI Reasoning systems can help answer the harder questions:

  • Why is this happening?  
  • What are the tradeoffs? 
  • What should we actually do next… Given everything we know about this asset, this site, and this moment?  

The Verdantix report notes that AI-driven planning tools have already reduced planning cycle times by 30 to 40 percent in real deployments. For UptimeAI customers, applying AI reasoning agents to close the decision gap in some of their most expert-intensive operations challenges has grown EBITDA margins by an average of 2-5%.  

Operational agility meets human resiliency 

In addition to the initial margin uplift realized when the insight-to-execution gap tightens, there are also some decidedly human benefits. When people aren’t buried in alert triage and manual data reconciliation, they think differently. They have space to use their judgment, to dig into their hunches. They start asking better questions, catch patterns earlier, and make the kinds of calls that only come from experience, and over time, they make them faster. And the AI learning from them gets smarter as a result.  

Operational resilience goes beyond building a better dashboard. It’s also about cultivating a team that’s operating at a higher level because the noise has been cleared out of their way. 

Make better decisions, faster, with the team you’ve got 

The Verdantix findings confirm what we’re seeing in our customers throughout the asset-intensive industries. The urgency is real, because the opportunity is real. And the organizations that close this gap first won’t just perform better today. They’re building a moat of competitive advantage that will make them structurally harder to catch for years to come. 

The State Of AI In Chemicals: Insights from CRU Nitrogen + Syngas Conference 2026 

By Cody Berra, CMRP, Senior Solutions Consultant, UptimeAI

Three pressures, one operational question — and why agentic operations
is becoming a C-suite conversation.

Ten years ago, the nitrogen playbook was relatively straightforward. Success came down to driving cost down, keeping rates high, and executing turnarounds cleanly. That model held up in a more stable market, where variability existed but could largely be managed within known bounds.

After three days at CRU Nitrogen and Syngas 2026 in Dallas, it is clear the above operating model is no longer sufficient for what is coming next. The kickoff session from Justin RackleffPrincipal Analyst, Fertilizer at CRU put a fine point on the shift. Roughly one third of global fertilizer supply has been disrupted by geopolitical conflict, forcing trade flows to be rerouted in real time rather than optimized over longer horizons. Prices are reacting quickly to these supply shocks, and even producers in the United States, despite their structural cost advantage, are feeling the impact of a more interconnected and volatile global market.

At the same time, the expectations placed on producers are increasing. There is a push to grow nitrogen and syngas output while also lowering carbon intensity, which introduces a layer of operational complexity that cannot be solved through capital investment alone. These pressures are not temporary dislocations that will normalize in a few quarters. Volatility, decarbonization, and the ongoing loss of experienced operators and engineers are becoming defining characteristics of how plants are run. 

Taken together, this shifts the conversation from strategy to execution. It is no longer just about where you sit on the cost curve, but how effectively you can respond as conditions change around you. That reality ultimately comes down to a single operational question, and the producers who can answer it consistently will separate themselves over the next decade.How quickly can your plant turn an alarm into the right action?

The Cheapest-Producer Thesis Is Breaking

For a long time, the answer to who wins in nitrogen production was simple. The lowest cost producer had the advantage, and over time that advantage compounded. In a more stable environment, that logic held up. 

What is changing is not that cost no longer matters, but that it is no longer enough on its own. In a market defined by volatility and tighter operating constraints, the ability to make the right decision at the right time is starting to matter just as much as the underlying cost position. Decision velocity, more than cents per ton, is beginning to shape outcomes. 

That showed up repeatedly in conversations throughout the week. The same underlying issue, something like early-stage fouling or a subtle process imbalance, can play out very differently depending on when and how it is addressed. When it is identified early, it becomes a manageable event that can be worked into a planned intervention. When it is missed or misdiagnosed, it has a way of escalating into an unplanned outage with meaningful financial impact. 

The difference is not the asset or even the condition itself. It is the speed and quality of the response. In a more volatile market, that gap becomes more expensive, because there is less room to absorb inefficiency or downtime. Small delays in recognizing what is happening, or uncertainty in what to do next, quickly translate into lost margin. 

This is where the traditional cost curve starts to break down. Cost will always be a factor, but it does not protect against slow or incorrect decisions. The producers who are pulling ahead are the ones who can consistently turn plant signals into the right operational action faster than the market can penalize a mistake. Over time, that capability compounds in the same way cost leadership once did. 

Decarbonization is now an operations problem, not just a capital one

The decarbonization conversation in nitrogen has historically been framed as a capital problem. Most of the focus has gone toward large investments like CCUS, blue ammonia, or electrolyzer-driven hydrogen. Those initiatives still matter and will continue to shape the long term direction of the industry. 

What came through more clearly this week is that carbon intensity is also a function of how well plants are run today. It is not just about what gets built next, but how close existing units operate to their intended design. 

A reformer running a few percentage points off target is not only leaving margin on the table, it is also increasing the carbon intensity of every ton produced. The same is true for fouling that goes undetected, heat transfer inefficiencies that persist longer than they should, or operating conditions that drift without being corrected. These are not just reliability or performance issues anymore, they are directly tied to emissions.


Decision velocity is now an ESG variable, highlighting the need for scaling expertise. 

In practice, that means decarbonization is no longer separate from day to day operations. Every delayed cleaning, every unplanned shutdown, and every period of suboptimal operation carries a carbon consequence that has to be accounted for elsewhere. Over time, that adds up in ways that are difficult to offset through capital projects alone. 

For producers investing in blue ammonia or low carbon hydrogen, this becomes even more important. Those projects are often modeled assuming a stable and well understood operating baseline. When the underlying unit is not consistently running close to design intent, that baseline becomes less reliable and the economics become harder to defend. 

The expertise crisis is already in the math

The workforce challenge came up in a more direct way than in prior years. It is no longer a future concern, it is already affecting day to day operations. When an experienced operator or specialist leaves, the impact goes beyond headcount. It is the loss of unit specific judgment built over decades, including how to interpret weak signals, recognize early deviations, and respond under non ideal conditions. That kind of knowledge is difficult to document and even harder to transfer in a way that is immediately useful. 

At the same time, the next generation is stepping into a more complex environment. Plants are dealing with more variability, tighter carbon constraints, and new technologies layered onto existing assets, all with less institutional knowledge to rely on. 

This is where the earlier pressures start to converge. Volatility increases the cost of slow decisions, decarbonization raises the stakes on how precisely plants are run, and the loss of experience reduces the amount of expert judgment available at any given moment. 


That gap is difficult to close through hiring alone. The constraint is no longer access to data, it is access to consistent, scalable decision making. 
And this decision-velocity gap, can only be resolved  by encoding expertise into the system itself.

 

Agentic operations is what closes the gap

Taken together, these pressures point to a shift in how plants need to operate. The question is no longer whether more data or better visibility is needed. Most sites already have that. The challenge is turning that information into the right decision quickly and consistently.

This is where the conversation around agentic operations is starting to take shape. Not as autonomous control, and not as another layer of dashboards, but as a way to close the gap between detecting an issue and knowing what to do about it. 

At most sites today, the real bottleneck is expert decision capacity. The ability to interpret signals in context, connect them to likely failure modes, and determine the right course of action is still heavily dependent on a small number of experienced individuals. As those individuals become harder to scale, the need to support that decision process in a more systematic way becomes more obvious. 

That is the role these systems are beginning to play. They are not replacing operators or engineers, but helping extend their judgment across more assets, more conditions, and more moments in time than would otherwise be possible. 

Three major shifts where agentic operations can empower senior leaders

  • From cost optimization to decision optimization. Cost will always matter. In this market, the producers who pull ahead will be the ones turning signals into action faster than their peers — at every site, every shift, every loop. 
  • From workforce planning to expertise institutionalization. You cannot hire your way out of the retirement math. You can encode the judgment of your best operators and engineers into systems that operate at machine speed across the fleet. 
  • From CCUS-as-investment to operations-as-decarbonization-lever. Carbon intensity now lives in how well today’s plants run, not just in tomorrow’s projects. The producers treating operational excellence as ESG strategy will compound an advantage the producers treating it as cost will miss.

We do not know what next year’s CRU will bring, but the conversation is already moving from “is this real?” to “how fast can we deploy?” That is a healthy shift for an industry that historically does not move first. Contact us to uncover how UptimeAI reasoning agents turn expertise into real-time, scalable advantage.