Infrastructure for a Moving Target | Part 1 of 3

AI Inference Is Moving Into Production. Enterprise Infrastructure Has to Move With It.

By Brad Alexander

Chief Technology Officer, DartPoints

In this series

This is Part 1 of Infrastructure for a Moving Target, a three-part series on how enterprise leaders can plan for placement, predictability, and control as workload requirements change.

Key takeaway

Production inference changes infrastructure planning because the workload is closer to the business. Latency, availability, consistency, connectivity, and placement become just as important as compute and storage.

Enterprise infrastructure used to move on a more predictable clock. You had refresh cycles, capacity
plans, cloud migration plans, and a set of assumptions that usually held long enough to make a decision.

That is not the environment most technology leaders are working in now.

The pace of change has picked up. New workload patterns are creating different pressure on power,
density, connectivity, and delivery timelines. The old question was often about how much compute and
storage the enterprise needed. That question still matters, but it is no longer enough.

For production inference, the better question is where the workload should run and how predictably it can
perform.

That shift sounds simple, but it changes the infrastructure conversation in a very practical way.

The workload is moving into the business

Training and inference are not the same infrastructure problem.

A lot of the early conversation around artificial intelligence focused on training. Training can often be centralized. It can be batched. It can happen in a place where distance from the user is not the biggest concern. If the job takes longer or runs in a large cluster somewhere else, the business may still be able to operate.

Inference is different.

Inference is where models start showing up in production environments that matter to the business. It may support a customer interaction, a workflow decision, a clinical process, a fraud model, a field
operation, or an internal productivity tool. Once that happens, latency, availability, and consistency become part of the equation.

That is when the conversation changes.

It is not just about having enough capacity. It is about whether the infrastructure is in the right place,
connected the right way, and able to support the experience the business expects.

Placement matters more than people think

When a workload is latency sensitive, geography starts to matter. When data needs to move between applications, users, and platforms, connectivity starts to matter. When a process becomes part of daily operations, availability starts to matter.

That does not mean every workload needs to sit five miles from every user. That would be too simple, and in many cases, too expensive.

It means enterprises need to be more intentional about placement.

Some workloads belong in a hyperscale environment because they need the tooling, ecosystem, and broad scale that comes with it. Some belong on premises because the business needs direct control. Some belong in a regional data center because they benefit from proximity, predictable performance, data control, and a better path to capacity. Some will need to connect across all of those environments.

The point is not to pick a single winner. The point is to stop treating all workloads the same.

For inference, success comes down to placement and predictability, not just scale.

Regional does not mean isolated

There is a misconception that regional infrastructure is only about being smaller or closer. That misses the real value.

A regional data center needs to help reduce latency where that matters, but it also needs to provide access into the larger ecosystem. That can mean connecting regional markets into major carrier hotels, internet exchanges, cloud onramps, and Tier 1 locations. When that is done well, an enterprise can
support local or regional performance requirements while still reaching the broader platforms it depends on.

That is an important middle ground.

A business may not need to run every workload in a primary market. It may not want the cost profile or power constraints that come with that. But it also cannot afford to strand workloads in places that are hard to connect, hard to operate, or hard to scale.

The right architecture keeps the business close to the users, close to the data, or close to the platform when it needs to be, without painting the organization into a corner.

Capacity is now part of strategy

The hard part of infrastructure is often where the plan meets the real world.

A strategy can look good in a meeting room, but it still has to survive power availability, equipment lead times, network delivery, facility readiness, and staffing realities. Those constraints are becoming more visible because the market is moving faster than traditional planning cycles.

For providers, that means capacity has to be planned well before the customer asks for it. For enterprises, it means waiting until the need becomes urgent can limit the options very quickly.

That is one of the biggest changes I see. The teams that handle this well are not necessarily the ones with the biggest plans. They are the ones that start earlier, understand the constraints, and build enough flexibility into the plan to adjust as conditions change.

The mistake is assuming capacity will be there when the business finally decides it needs it.

“The hard part of infrastructure is often where the plan meets the real world.”

Flexibility is not the same as uncertainty

Some leaders hear stay flexible and think it means delaying decisions. That is not what I mean.

Flexibility should be designed into the infrastructure model. It means knowing which workloads need low latency, which need predictable cost, which need direct control, and which need large-scale cloud  tooling. It means building a path where the business can move in stages without having to re-architect everything each time conditions change.

That matters because most organizations do not have perfect visibility into what their technology environment will need 18 to 24 months from now. The target is moving. The answer is not to pretend it is standing still.

The answer is to create an operating model that gives the business room to adjust.

What I would tell enterprise leaders

Do not over-optimize for a fixed end state.

That is probably the most important advice I would give any CIO or CTO looking at production inference workloads. The end state will change. The requirements will change. The vendor landscape will change. The power and capacity picture may change too.

What should not change is the need for predictability, execution, and control.

Start with the workload. Understand the latency profile. Understand the data. Understand the availability requirements. Understand the cost model. Then decide where it belongs.

Some of that will land in hyperscale. Some may stay on premises. Some will fit better in a regional colocation or hybrid environment. Most enterprises will end up with a mix.

That is not a failure of strategy. That is the strategy.

The goal is to put the right workload in the right place, then keep enough flexibility in the model to stay effective as things evolve.

Continue the series | Part 2
Why regulated industries need data control before speed.

Not sure where production workloads should run?

DartPoints can help evaluate workload placement, latency needs, connectivity, and capacity planning across regional colocation, cloud, and hybrid environments.
Guidelines

File info
Format
File size
Resolution