AI-OPS · Path 4: Build and operate
LLMOps for Connected, Disconnected and Air-Gapped Environments
Three days after which a model update arrives safely even when no cable leads to the outside world.
- Duration
- 3 days
- Group
- up to 12 people
- Format
- In-house · In the lab
- Language
- German or English
- Price
- €23,700 flat, plus VAT
Who this track is for
Platform and operations teams responsible for an LLM platform in environments with restricted or no internet connectivity: public authorities, critical infrastructure, defense, manufacturing with segregated networks. Also for teams whose platform is connected today but who need a credible separation scenario.
Who it is not for
Not for teams that do not yet run or plan a platform. They build the foundation first in Sovereign AI Platform Engineering (AI-PLATFORM).
Starting situation
Your AI platform is supposed to run in an environment where the usual MLOps recipes fail: no access to Hugging Face, no container registry on the internet, no telemetry backend in the cloud. Models, containers and signatures must pass a controlled gateway, updates must not break anything, and a rollback has to work without vendor support from outside.
This exists after the track
- A documented artifact flow for models, containers, prompts and configurations with versioning and a signature chain
- A dev-test-prod promotion process with defined gates, including an evaluation gate before every model change
- An offline update concept: transfer package, gateway process, signature verification and staging without internet access, played through once in full in the lab
- An incident and rollback runbook for model regression, compromised artifacts and platform outage, tested in an exercise
- A monitoring concept that works without a cloud backend and still makes model quality visible
Prerequisites
Operational experience with containers and a basic understanding of CI/CD. Having attended AI-PLATFORM or running an existing inference environment in-house is ideal.
Preparation before the track
You sketch your target environment on one page in advance: network zones, permitted transitions, today's update paths and the one incident you fear most at night. These sketches become the group's case studies on day one.
What's included
- 3 training days delivered by Dino Bordonaro
- A lab environment with a deliberately isolated network zone per team where the offline exercises take place for real
- Templates for the transfer package manifest, promotion gates and the rollback runbook
- Certificates of attendance and audit-proof training documentation
- A 60-minute follow-up with the team 4 weeks after the course
Agenda
Day 1
Three operating models, three truths
Connected, intermittently disconnected, permanently disconnected: what really changes per model in update frequency, monitoring and supportability. Honest framing of Azure Local Disconnected Operations: access-restricted, subject to an eligibility check, not a freely orderable product.
Everything is an artifact
Models, containers, prompts, configurations and evaluation data as versioned, signed artifacts. A concrete question per artifact type: where is it created, who signs it, what does the signature prove?
Promotion gates that actually stop things
The path from dev through test to prod with gates that are armed: evaluation suite, signature verification, four-eyes approval. Case discussion: which gate would have stopped your organization's last bad model change?
Exercise: promotion pipeline in the lab
Each team builds the pipeline: a model artifact passes dev and test, a deliberately degraded model must get stuck at the evaluation gate. If it does not, the gate gets sharpened.
Day result: documented promotion process
Each team has a promotion process with defined gates as a document and as a running pipeline, proven by one model that passed and one that was stopped.
Day 2
Anatomy of a transfer package
What passes through the gateway: model weights, container images, signatures, manifests, evaluation data. Building a manifest that makes completeness and integrity verifiable before anything gets installed.
The gateway as a process, not a USB stick
Roles, verification steps and logging at the transition into the disconnected zone. Which checks run outside, which inside, and who owns the approval? Cross-checked against the network sketches you brought.
Exercise: offline update under real conditions
Each team's lab zone is cut off from the internet. A new model and its container are built into a transfer package, brought through the gateway, verified against the manifest and activated in staging. A tampered package is in circulation and must fail signature verification.
Debrief: where it jammed
Review of the exercise along the logs: which check was missing, which was redundant, how long did the run take? The findings flow straight into your own update concept.
Day result: completed offline update with protocol
Each team has performed a full update without internet access and demonstrably rejected the tampered package. The gateway protocol documents every step.
Day 3
Monitoring without a cloud backend
Making model quality, drift and platform health visible when no SaaS dashboard is allowed. Building the metrics pipeline in the disconnected zone, including the question of which evaluations may leave periodically and which never.
Incident scenarios for LLM platforms
Model regression after an update, suspected compromised artifact, GPU failure, a wave of prompt injection: for each scenario the runbook records detection path, immediate action and communication chain.
Exercise: rollback under time pressure
The model installed the day before shows a massive regression in the exercise. Each team rolls back to the last known good state by runbook, measures the time and proves via evaluation that the old state is restored.
Your 90-day plan for real operations
Team prioritization: which three gaps between lab exercise and your own environment get closed first, who owns the signature chain in your organization, when does the first in-house gateway test run take place.
Day result: tested incident and rollback runbook
Each team has a runbook that survived a real rollback exercise, including measured recovery time and a documented 90-day plan for its own environment.
Exercises and lab share
About 60 percent of the time is exercises: each team operates a genuinely isolated lab zone, builds transfer packages, rejects a tampered package and performs a complete rollback against the clock.
Platforms
Delivered on our lab environment with separated network zones, on request in-house on your own infrastructure if access and approvals are in place. We treat Azure Local Disconnected Operations honestly as an access-restricted offering with an eligibility check, and all platform features apply in the version available at the time. BORDONARO is an independent partner, not a Microsoft subsidiary.
Transfer evidence
Three documented proofs: the pipeline that stops a bad model, the gateway protocol of the offline update with the rejected tampered package and the timed rollback exercise. All three are handed to the participants as documentation.
Artifacts you take home
- Dev-test-prod promotion process as a document and pipeline template
- Transfer package manifest and gateway protocol as templates
- Incident and rollback runbook, tested in the exercise
- Monitoring concept for environments without a cloud backend
- 90-day plan for the transfer into your own environment
Optional extensions
- Azure Arc and Azure Local as a Hybrid Control Plane (AI-ARC) for the connected versus disconnected decision picture
- Evaluation, Observability and Guardrails (AI-EVAL) to deepen the evaluation gates
- A guided first gateway test run in your environment as a separate service
Boundaries
The course delivers processes, templates and rehearsed procedures, not a certification of your environment and not an implementation of your gateway. Build-out and operations in your organization are project services and are contracted separately.
Frequently asked questions
Will our environment be air-gapped and secure after the course?
No, and beware of anyone who promises that. The course teaches the processes and evidence with which you make a disconnected environment operable. Whether and how strictly your separation is implemented is decided by your security architecture and your regulators, not by a training.
Can this be done remotely?
The core of the course lives on the isolated lab zones and the physical gateway exercise, so we deliver it in person only: in the lab or in-house at your site. The follow-up after 4 weeks takes place remotely.
We are fully connected today. Is the course still worth it?
Yes, if a separation scenario could become real, for example through regulatory conditions, customers in critical infrastructure or certification goals. The promotion process and signature chain from day one improve connected environments immediately, the rest is your insurance.