Skip to content

AI-DATA · Path 4: Build and operate

AI-ready Data and Knowledge Architecture

RAG projects rarely fail at the model, they fail at sources without owners and permissions without a concept.

Duration
2 days
Group
up to 12 people
Format
In-house · Remote
Language
German or English
Price
€15,800 flat, plus VAT

Who this track is for

Architects, data engineers and developers tasked with making company knowledge usable for assistants. Also for technical project leads who want to know which sources can carry a RAG initiative before it starts.

Who it is not for

Not for teams that already have a solid source map and want to build the prototype. They start directly in Enterprise RAG Engineering (AI-RAG). Pure policy questions belong in AI-GOV.

Starting situation

Your knowledge lives in SharePoint, file servers, wikis, ticket systems and mailboxes, often in three versions and without a clear owner. An assistant answering on top of that estate cites outdated policies and surfaces documents the person asking should never see. Before a single line of ingestion code is written, you need a map, ownership and a permission concept.

This exists after the track

  • A knowledge map of your real source landscape, developed in anonymized form, rated by freshness, quality and access structure
  • An ownership matrix: a named business owner and a defined update process for every prioritized source
  • A prioritized source list for the first RAG iteration with documented exclusion criteria
  • A retrieval requirements profile: which question types, document types and metadata the system must serve
  • A draft permission concept that carries source ACLs all the way to the answer instead of losing them at indexing time

Prerequisites

Familiarity with your own system landscape and a basic understanding of data structures. Programming skills are not required, architectural thinking is.

Preparation before the track

You receive a source profile template in advance: per participating organization the 10 most important knowledge sources with system type, estimated volume, access model and last known cleanup date. The anonymized profiles are the working material for both days.

What's included

  • 2 training days with Dino Bordonaro, on site or remote
  • Reusable templates for knowledge map, ownership matrix and source assessment
  • A reference permission concept with a decision tree for document-level security
  • An anonymized evaluation of the source profiles as a situation report
  • Certificates of attendance and training documentation for your compliance archive
  • A 60-minute remote office hour 4 weeks after the training to review your map

Agenda

Day 1

09:00

The situation report: your sources on the table

Review of the anonymized source profiles. First sorting question: which of these sources would an assistant cite today, and whom would you trust with that answer?

10:30

Why indexing without a map gets expensive

Anatomy of failed RAG projects: stale duplicates, orphaned drives, ACL loss during crawling. Live demo on the Sovereign Assistant in our lab: the same question against curated and uncurated sources.

13:00

Knowledge map workshop

Each organization maps its sources along three axes: business value, data quality, access structure. The result is a four-quadrant map with clear candidates and clear exclusions.

15:00

Ownership or data graveyard

Exercise on your own maps: who owns each source, who maintains it, what happens when two versions contradict each other? The ownership matrix is built on your own case.

16:30

Day result: map and matrix stand

Each organization presents its knowledge map with ownership matrix in 5 minutes. The group's test question: is every top source assigned to a name, and is every exclusion justified?

Day 2

09:00

Prioritization: cutting the first iteration

The map becomes an ordered list: which 3-5 sources carry the first RAG iteration? The criteria are answer value, data readiness and permission clarity, not political visibility.

10:30

Retrieval requirements instead of a wish list

Real question types from the participants become a requirements profile: fact questions, process questions, comparison questions. Which metadata, structures and freshness guarantees does each class need?

13:00

Permissions: from source to answer

Workshop on the reference concept: how do ACLs travel from SharePoint and file servers into the index, and how are they enforced on every query? Decision tree for document-level security applied to your own access model.

15:00

The maintenance process: staying current instead of cleaning up once

Exercise: the update path is defined for each prioritized source. Who deletes, who versions, how does the index learn about it? Without this process, every map goes stale within months.

16:15

Day result: the handover-ready package

Each organization files its map, ownership matrix, prioritized source list and draft permission concept as one package. Peer review against a checklist: could a RAG team start with this on Monday?

Exercises and lab share

Around half the time is workshop work on your own anonymized source landscape. Live demonstrations on the Sovereign Assistant in our lab environment make curation and permission effects visible.

Platforms

Workshop work uses anonymized profiles of your real sources, no customer data leaves your organization. Demos run on our lab environment with synthetic data, platform examples in the respective available version.

Transfer evidence

The final review on Day 2 checks every package against a documented checklist: ownership named, prioritization justified, permission path described. Attendance and results are documented in an audit-proof way.

Artifacts you take home

  • Knowledge map of your own source landscape as an editable document
  • Ownership matrix with a maintenance process per prioritized source
  • Prioritized source list with exclusion rationale for the first RAG iteration
  • Draft permission concept with a decision tree for document-level security
  • Retrieval requirements profile as the input document for AI-RAG

Optional extensions

  • Enterprise RAG Engineering (AI-RAG) as the direct build follow-up based on the package
  • Evaluation, Observability and Guardrails (AI-EVAL) for the quality assurance of the future system
  • A Sovereign Assistant PoC on our lab environment with your prioritized sources as a project

Boundaries

This training produces concepts and decision material for your case. It replaces neither data migration nor system rollout nor an implementation engagement, for those we talk about a project.

Frequently asked questions

Our data landscape is a mess. Isn't this training premature?

That is exactly when it is right. The goal is not a tidy landscape but a justified selection of the 3-5 sources that can carry a first iteration. Cleaning up everything before the first project is the most common way to never start.

Do we have to bring real company data?

No. The work is done on anonymized source profiles: system type, volume, access model, quality assessment. That is enough for sound decisions, and no content leaves your organization.

We have not decided on a RAG project yet. Is this still worth it?

Yes, because the package of map, ownership and permission concept is also the basis for deciding whether and where a project pays off. If the business case question comes first, AI-VALUE is the better entry point.