Skip to main content
🚀 Taking AI from prototype to production? Find the architecture, GPU, security and governance gaps before they become incidents. Get a Production AI Readiness Assessment
Luca Berton pointing at his LUCA BERTON speaker card on the AI Tech Summit sponsor wall in Skopje
Conferences

AI Tech Summit Skopje 2026: From AI Demo to Production

My keynote at AI Tech Summit Filip Avramchev 2026 in Skopje, the sessions I photographed, and a practical checklist for taking AI from demo to production.

LB
Luca Berton
¡ 11 min read

On 22 and 23 September 2026 I was in Skopje, North Macedonia, as a speaker at AI Tech Summit “Filip Avramchev” 2026. The summit took place at the National Opera and Ballet. I gave a keynote on the second morning: “From AI Demo to Production: Building AI Platforms That Actually Scale”. This post covers the summit, what my session was about, the other sessions I photographed over the two days, and a checklist that expands on the themes of my talk.

The main auditorium of the National Opera and Ballet in Skopje, seen from the back rows, with the illuminated ai tech summit sign above the stage

The Main Stage in the opera’s main auditorium on the morning of Day 1.

About the summit and its founder

The official site describes AI Tech Summit “Filip Avramchev” as “a leading regional AI summit based in Skopje, North Macedonia, dedicated to advancing the understanding and responsible adoption of artificial intelligence across industries and society”. It says the summit aims to strengthen the AI ecosystem across Southeast Europe and beyond.

The name in quotation marks is a dedication. According to the site, the summit was founded by Filip Avramchev, “with a clear vision: to build a trusted platform where regional talent and innovation meet the global AI ecosystem”. The site says that although he has passed away, “his vision lives on”, and it calls the 2026 edition “the fifth milestone of this journey”. The Recursive’s event listing also says the edition honours “the late Filip Avramchev”. The speaker list names Viktor Sushelski of Cyborg Code Syndicate as the summit’s founder and organiser.

The agenda was split into four areas: the Main Stage, AI Labs, the Hackathon and an Expo. The AI Student Hackathon brought together high school and university teams of up to four. It started on 18 September and gave the teams five days to build their solutions with OpenAI tools and help from mentors. The Top 3 teams then presented in a Grand Final on the Main Stage on 23 September. Main Stage speakers came from OpenAI, NVIDIA, Microsoft, Waymo, Zalando, KPMG, IDEO and EMOTIV, among others, plus companies from the region.

Luca Berton pointing at the LUCA BERTON speaker card on the AI Tech Summit sponsor wall, surrounded by sponsor and partner logos

My speaker card on the sponsor wall in the foyer. The same card also appeared on the logo board above the Main Stage.

My keynote: From AI Demo to Production

The agenda listed my session as a 30-minute keynote on the Main Stage on Day 2 (Wednesday 23 September), in the “MLOps & Platforms” track, right after Dr. Irena Bojarovska’s Zalando keynote. I was billed as Founder & CEO of Open Empower and an AI platform engineering educator and advisor.

The published abstract sets out what the session was about:

Moving an AI application from an impressive demo to a reliable production system is where the real engineering begins. In this session, Luca Berton will explore how platform engineering can provide the foundation for production AI—from containerized workloads and Kubernetes to GPU orchestration, observability, security, governance and cost control. Attendees will learn the practical architectural decisions that help teams scale AI responsibly while giving developers a faster, safer path from experimentation to production.

In short, the argument is that a model that works in a demo is not yet a product. What turns it into one is a platform underneath: Kubernetes and containers as the runtime, GPUs scheduled and shared deliberately, telemetry, security and policy built in, and costs that someone can see and own. The point of all that is to give developers a faster, safer way from a notebook to a production endpoint. I don’t have a photo of myself on stage, so this post uses the abstract rather than slides. The checklist further down expands each of those themes into concrete steps.

Day 1: agents, human-centred AI and sovereign blueprints

Panel: Beyond Copilots

Around midday on Day 1 I photographed the panel “Beyond Copilots: Agents and the Future of Software Engineering”. The slide named Anelija Mitrova (Program Coordinator & Moderator, AI Tech Summit Filip Avramchev | Youth Alliance – Krusevo) as moderator. The panellists were Alistair Greenwood (Founder, OmniforceAI), James Pelton (Developer Relations, Zencoder), Natasha Milenkovska Talevska (Director of Data, IWConnect) and Miodrag Stojanov (Head Of Engineering, Semos Cloud).

The Beyond Copilots: Agents and the Future of Software Engineering panel on stage, with a slide listing the moderator and four panellists above them

The “Beyond Copilots” panel line-up slide, with the panellists seated on stage.

IDEO: user research isn’t dead

In the afternoon I photographed the IDEO session. The agenda lists it as Angela Kochoska (Senior Data Science Lead, IDEO) with the keynote “The Value and Intersections of Artificial and Human Intelligences in Innovation Work”. An early slide headed “IDEO · AI & Emerging Tech” showed three projects: Ethiqly, described as “an AI-native edtech startup for writing assistance”; “Digital Twin networking at MIT Technology Review”; and “GenAI & View-Masters for futuring”. Later slides in the same session read “That’s so agentic”, “Desirability → Utility”, “From LLMs to AI-generated Users (today)”, a Waymo “Digital Twins” example, “Data galore! At lightning speeds!” and a cycle of Understand, Design, and Test + Iterate under the heading “The Anxiety Gap”.

On the Main Stage, a speaker in a red dress next to a slide reading User research isn't dead, after all..., with a billboard photo that says Listen to people. Don't replace them.

“User research isn’t dead, after all…”, illustrated with a billboard reading “Listen to people. Don’t replace them.”

Sovereign by Design: three blueprints

In the AI Labs programme, in front of the AI Student Hackathon backdrop (“Build. Learn. Present. Win.”), Gjorgji Dimitrov presented “Sovereign by Design: Three blueprints for AI that can’t leave the building”. His title slide listed him as Founding CTO of GAIA Technology Systems and CEO of Qualimetrix. Its subtitle was “Where your data actually goes - and what each answer costs to build, run and defend.”

I photographed the first two blueprints:

  • Blueprint 1: Managed API, public model. “Your application calls a hosted frontier model. The provider runs everything behind the API.” The slide said you control prompts, retrieval, retention settings and what you choose to send. You trust the provider’s isolation, its contractual no-training terms and its subprocessors. It called this the right answer for internal productivity, public content, prototypes and “anything non-personal”, and added: “This is the correct answer far more often than sovereignty rhetoric admits. Fastest path to value, lowest cost, best model quality.”
  • Blueprint 2: Private endpoints in your own tenant. “The model deployment lives inside your cloud tenant, in your region, reachable only over private networking.” The architecture box was labelled “YOUR TENANT ¡ YOUR REGION ¡ NO PUBLIC EGRESS” and contained your app, a private endpoint, the model deployment, a permission-filtered retrieval index and an audit log in your own storage. The slide gave it as the right answer for “Personal or financial data, DPA-bound work, most regulated EU delivery”.

A speaker in front of a projected slide titled Private endpoints in your own tenant, with an architecture diagram marked your tenant, your region, no public egress

Blueprint 2 of 3: “where a security questionnaire stops being an argument and becomes a diagram”.

My take: this is a useful way to frame sovereignty because it starts from the data, not from the slogan. In my own client work, the second blueprint is also where platform engineering takes on most of the load: private networking, identity, retention and audit logs are all platform concerns. My notes on digital sovereignty in Europe go into the wider strategy.

A community partner: AI NOW

On the morning of Day 1, I recorded a short interview for my show with the founder of AI NOW, whose site describes it as the Artificial Intelligence Association in North Macedonia. The founder said AI NOW is a partner community of the summit, and that it focuses on educating people and organisations about AI in the simplest possible terms, because “not everyone is a tech guy”. It also runs AI Balkans, a hub for AI news and tools from the region. The founder described its Edu AI Now tool as a platform you download to learn prompting, AI literacy and agents. It is “offline first”, has no trackers, and runs on a desktop without an internet connection. According to the founder, communities, private schools and universities use it, with about 70 students so far.

Day 2: forecasting, agentic workflows and the human mandate

The Accelerated Scientist (Zalando)

The first Day 2 keynote on the agenda was “The Accelerated Scientist: From Forecast Accuracy to Commercial Impact” by Dr. Irena Bojarovska (Applied Scientist, Zalando SE). According to the agenda abstract, the talk followed Zalando’s forecasting work: what time-series LLMs can offer, including “the potential to substitute hundreds of specialized models with a single global one”, the practical problems of applying them to real retail data, and why forecast stability matters as much as accuracy for commercial steering.

The slides I photographed matched that outline: “The Maslow Pyramid of Time-Series Forecasting”, “The Promise: One Model to Forecast It All” and “Chronos-2: The Evolution from Univariate to Universal Time Series Forecasting”. That slide described “zero-shot support for univariate, multivariate, and covariate-informed forecasting tasks” and showed the architecture from input through tokenisation and a transformer stack with time and group attention. Later slides read “Accuracy alone cannot build operational trust” and “How to measure forecast stability?”.

Dr. Irena Bojarovska on the Main Stage next to a slide titled Chronos-2: The Evolution from Univariate to Universal Time Series Forecasting, with an architecture diagram

The Chronos-2 slide: zero-shot support for univariate, multivariate and covariate-informed forecasting.

Near the end of the talk, a slide titled “Building in 1 Day: The Agentic Workflow in Practice” listed five steps:

  1. Define the “Ingredients”: supply the necessary git repositories, .md files, historical data and so on.
  2. Formulate the Agentic Prompt: convert the goal into a model instruction, and be as clear and as detailed as possible.
  3. Refine strategy and assumptions: answer any questions and verify the agent’s assumptions.
  4. Have a coffee: start the implementation (the rest of the line was hidden behind the speaker).
  5. Test: use the produced solution and check its usability (partly hidden too).

A note in the corner asked: “But maybe prompting is dead in 6 months?”, attributed to Andrew Ng.

The Building in 1 Day: The Agentic Workflow in Practice slide with five numbered steps, next to the speaker on stage

Five steps from ingredients to test: the agentic workflow behind a one-day prototype.

My take: steps 1 and 3 are platform problems in disguise. An agent can only build in a day if the repositories, docs and data it needs are discoverable and accessible under the right permissions. That is exactly what a good internal platform provides.

The Human Mandate

Later that morning I photographed a slide branded CyborgCEO in what the agenda lists as Darren Goonawardana’s (CEO & Co-Founder, Collectiv) keynote “The Human Mandate”. The slide read: “When customers can get the service autonomously, human contact can be the feature. Not just the fallback.”

A packed opera auditorium watching a speaker on the Main Stage next to a slide reading human contact can be the feature

A full auditorium for “The Human Mandate” on Day 2.

From AI demo to production: a practical checklist

This section expands on the themes in my keynote abstract. It reflects how I work with teams, not a transcript of the talk.

1. Containerise the workload and run it on Kubernetes. Package the model server, the application and its dependencies as images with pinned versions. Deploy them declaratively, through GitOps if you can, so that the demo and production environments differ only in configuration. What Breaks First When AI Moves from PoC to Production lists the assumptions that usually fail first.

2. Treat GPUs as a shared, scheduled platform resource. Decide up front how teams get GPU capacity: dedicated nodes, MIG partitions, time-slicing or quotas per tenant. Make that decision explicit. My multi-tenant GPU orchestration on OpenShift AI post and the KubeCon Europe 2026 talk recap cover how this works in practice. GPU Sharing on Kubernetes: MIG vs MPS vs Time-Slicing compares the options.

3. Serve models behind a proper serving layer with SLOs. Pick a serving runtime (vLLM, Triton, NIM, or KServe on OpenShift AI), define latency SLOs such as time to first token, and benchmark against them before launch. See AI Model Serving on Kubernetes, Running LLMs on OpenShift AI, benchmarking vLLM with GuideLLM and autoscaling AI inference with HPA and KEDA.

4. Make it observable from day one. Collect GPU utilisation, token latency, throughput, error rates and cost per request in the same stack you use for everything else, and trace requests end to end. AI Observability on Kubernetes has the metrics and dashboards I start with.

5. Secure the path, not only the model. Use per-workload identity, least-privilege access to data and tools, private networking where the data requires it (the “Sovereign by Design” blueprint 2 above), and signed images. For agents, add approval gates, blast-radius limits and audit logs, as described in Guardrails for AI Agents in Production.

6. Turn governance into policy as code. Write down which models, registries, regions and data classes are allowed, and enforce those rules at admission time instead of relying on review meetings. Kyverno and policy-driven AI governance shows one way to do it on Kubernetes.

7. Put cost control in from the start. Tag every GPU workload with an owner and a cost centre, give teams showback, right-size and autoscale inference, and switch off idle capacity. FinOps for AI: Control GPU Costs has the playbook.

8. Give developers a golden path. Package all of the above as a template: a model service with CI, policies, dashboards and cost tags already wired in, available through an internal developer platform. That is where the “faster, safer path from experimentation to production” comes from. See Internal Developer Platforms Compared.

If your team is stuck between a promising demo and a production rollout, this is the kind of work I do in AI integration consulting.

Free 30-min Production AI consultation

Book Now