On 22 and 23 September 2026 I was in Skopje, North Macedonia, as a speaker at AI Tech Summit âFilip Avramchevâ 2026. The summit took place at the National Opera and Ballet. I gave a keynote on the second morning: âFrom AI Demo to Production: Building AI Platforms That Actually Scaleâ. This post covers the summit, what my session was about, the other sessions I photographed over the two days, and a checklist that expands on the themes of my talk.

The Main Stage in the operaâs main auditorium on the morning of Day 1.
About the summit and its founder
The official site describes AI Tech Summit âFilip Avramchevâ as âa leading regional AI summit based in Skopje, North Macedonia, dedicated to advancing the understanding and responsible adoption of artificial intelligence across industries and societyâ. It says the summit aims to strengthen the AI ecosystem across Southeast Europe and beyond.
The name in quotation marks is a dedication. According to the site, the summit was founded by Filip Avramchev, âwith a clear vision: to build a trusted platform where regional talent and innovation meet the global AI ecosystemâ. The site says that although he has passed away, âhis vision lives onâ, and it calls the 2026 edition âthe fifth milestone of this journeyâ. The Recursiveâs event listing also says the edition honours âthe late Filip Avramchevâ. The speaker list names Viktor Sushelski of Cyborg Code Syndicate as the summitâs founder and organiser.
The agenda was split into four areas: the Main Stage, AI Labs, the Hackathon and an Expo. The AI Student Hackathon brought together high school and university teams of up to four. It started on 18 September and gave the teams five days to build their solutions with OpenAI tools and help from mentors. The Top 3 teams then presented in a Grand Final on the Main Stage on 23 September. Main Stage speakers came from OpenAI, NVIDIA, Microsoft, Waymo, Zalando, KPMG, IDEO and EMOTIV, among others, plus companies from the region.

My speaker card on the sponsor wall in the foyer. The same card also appeared on the logo board above the Main Stage.
My keynote: From AI Demo to Production
The agenda listed my session as a 30-minute keynote on the Main Stage on Day 2 (Wednesday 23 September), in the âMLOps & Platformsâ track, right after Dr. Irena Bojarovskaâs Zalando keynote. I was billed as Founder & CEO of Open Empower and an AI platform engineering educator and advisor.
The published abstract sets out what the session was about:
Moving an AI application from an impressive demo to a reliable production system is where the real engineering begins. In this session, Luca Berton will explore how platform engineering can provide the foundation for production AIâfrom containerized workloads and Kubernetes to GPU orchestration, observability, security, governance and cost control. Attendees will learn the practical architectural decisions that help teams scale AI responsibly while giving developers a faster, safer path from experimentation to production.
In short, the argument is that a model that works in a demo is not yet a product. What turns it into one is a platform underneath: Kubernetes and containers as the runtime, GPUs scheduled and shared deliberately, telemetry, security and policy built in, and costs that someone can see and own. The point of all that is to give developers a faster, safer way from a notebook to a production endpoint. I donât have a photo of myself on stage, so this post uses the abstract rather than slides. The checklist further down expands each of those themes into concrete steps.
Day 1: agents, human-centred AI and sovereign blueprints
Panel: Beyond Copilots
Around midday on Day 1 I photographed the panel âBeyond Copilots: Agents and the Future of Software Engineeringâ. The slide named Anelija Mitrova (Program Coordinator & Moderator, AI Tech Summit Filip Avramchev | Youth Alliance â Krusevo) as moderator. The panellists were Alistair Greenwood (Founder, OmniforceAI), James Pelton (Developer Relations, Zencoder), Natasha Milenkovska Talevska (Director of Data, IWConnect) and Miodrag Stojanov (Head Of Engineering, Semos Cloud).

The âBeyond Copilotsâ panel line-up slide, with the panellists seated on stage.
IDEO: user research isnât dead
In the afternoon I photographed the IDEO session. The agenda lists it as Angela Kochoska (Senior Data Science Lead, IDEO) with the keynote âThe Value and Intersections of Artificial and Human Intelligences in Innovation Workâ. An early slide headed âIDEO ¡ AI & Emerging Techâ showed three projects: Ethiqly, described as âan AI-native edtech startup for writing assistanceâ; âDigital Twin networking at MIT Technology Reviewâ; and âGenAI & View-Masters for futuringâ. Later slides in the same session read âThatâs so agenticâ, âDesirability â Utilityâ, âFrom LLMs to AI-generated Users (today)â, a Waymo âDigital Twinsâ example, âData galore! At lightning speeds!â and a cycle of Understand, Design, and Test + Iterate under the heading âThe Anxiety Gapâ.

âUser research isnât dead, after allâŚâ, illustrated with a billboard reading âListen to people. Donât replace them.â
Sovereign by Design: three blueprints
In the AI Labs programme, in front of the AI Student Hackathon backdrop (âBuild. Learn. Present. Win.â), Gjorgji Dimitrov presented âSovereign by Design: Three blueprints for AI that canât leave the buildingâ. His title slide listed him as Founding CTO of GAIA Technology Systems and CEO of Qualimetrix. Its subtitle was âWhere your data actually goes - and what each answer costs to build, run and defend.â
I photographed the first two blueprints:
- Blueprint 1: Managed API, public model. âYour application calls a hosted frontier model. The provider runs everything behind the API.â The slide said you control prompts, retrieval, retention settings and what you choose to send. You trust the providerâs isolation, its contractual no-training terms and its subprocessors. It called this the right answer for internal productivity, public content, prototypes and âanything non-personalâ, and added: âThis is the correct answer far more often than sovereignty rhetoric admits. Fastest path to value, lowest cost, best model quality.â
- Blueprint 2: Private endpoints in your own tenant. âThe model deployment lives inside your cloud tenant, in your region, reachable only over private networking.â The architecture box was labelled âYOUR TENANT ¡ YOUR REGION ¡ NO PUBLIC EGRESSâ and contained your app, a private endpoint, the model deployment, a permission-filtered retrieval index and an audit log in your own storage. The slide gave it as the right answer for âPersonal or financial data, DPA-bound work, most regulated EU deliveryâ.

Blueprint 2 of 3: âwhere a security questionnaire stops being an argument and becomes a diagramâ.
My take: this is a useful way to frame sovereignty because it starts from the data, not from the slogan. In my own client work, the second blueprint is also where platform engineering takes on most of the load: private networking, identity, retention and audit logs are all platform concerns. My notes on digital sovereignty in Europe go into the wider strategy.
A community partner: AI NOW
On the morning of Day 1, I recorded a short interview for my show with the founder of AI NOW, whose site describes it as the Artificial Intelligence Association in North Macedonia. The founder said AI NOW is a partner community of the summit, and that it focuses on educating people and organisations about AI in the simplest possible terms, because ânot everyone is a tech guyâ. It also runs AI Balkans, a hub for AI news and tools from the region. The founder described its Edu AI Now tool as a platform you download to learn prompting, AI literacy and agents. It is âoffline firstâ, has no trackers, and runs on a desktop without an internet connection. According to the founder, communities, private schools and universities use it, with about 70 students so far.
Day 2: forecasting, agentic workflows and the human mandate
The Accelerated Scientist (Zalando)
The first Day 2 keynote on the agenda was âThe Accelerated Scientist: From Forecast Accuracy to Commercial Impactâ by Dr. Irena Bojarovska (Applied Scientist, Zalando SE). According to the agenda abstract, the talk followed Zalandoâs forecasting work: what time-series LLMs can offer, including âthe potential to substitute hundreds of specialized models with a single global oneâ, the practical problems of applying them to real retail data, and why forecast stability matters as much as accuracy for commercial steering.
The slides I photographed matched that outline: âThe Maslow Pyramid of Time-Series Forecastingâ, âThe Promise: One Model to Forecast It Allâ and âChronos-2: The Evolution from Univariate to Universal Time Series Forecastingâ. That slide described âzero-shot support for univariate, multivariate, and covariate-informed forecasting tasksâ and showed the architecture from input through tokenisation and a transformer stack with time and group attention. Later slides read âAccuracy alone cannot build operational trustâ and âHow to measure forecast stability?â.

The Chronos-2 slide: zero-shot support for univariate, multivariate and covariate-informed forecasting.
Near the end of the talk, a slide titled âBuilding in 1 Day: The Agentic Workflow in Practiceâ listed five steps:
- Define the âIngredientsâ: supply the necessary git repositories, .md files, historical data and so on.
- Formulate the Agentic Prompt: convert the goal into a model instruction, and be as clear and as detailed as possible.
- Refine strategy and assumptions: answer any questions and verify the agentâs assumptions.
- Have a coffee: start the implementation (the rest of the line was hidden behind the speaker).
- Test: use the produced solution and check its usability (partly hidden too).
A note in the corner asked: âBut maybe prompting is dead in 6 months?â, attributed to Andrew Ng.

Five steps from ingredients to test: the agentic workflow behind a one-day prototype.
My take: steps 1 and 3 are platform problems in disguise. An agent can only build in a day if the repositories, docs and data it needs are discoverable and accessible under the right permissions. That is exactly what a good internal platform provides.
The Human Mandate
Later that morning I photographed a slide branded CyborgCEO in what the agenda lists as Darren Goonawardanaâs (CEO & Co-Founder, Collectiv) keynote âThe Human Mandateâ. The slide read: âWhen customers can get the service autonomously, human contact can be the feature. Not just the fallback.â

A full auditorium for âThe Human Mandateâ on Day 2.
From AI demo to production: a practical checklist
This section expands on the themes in my keynote abstract. It reflects how I work with teams, not a transcript of the talk.
1. Containerise the workload and run it on Kubernetes. Package the model server, the application and its dependencies as images with pinned versions. Deploy them declaratively, through GitOps if you can, so that the demo and production environments differ only in configuration. What Breaks First When AI Moves from PoC to Production lists the assumptions that usually fail first.
2. Treat GPUs as a shared, scheduled platform resource. Decide up front how teams get GPU capacity: dedicated nodes, MIG partitions, time-slicing or quotas per tenant. Make that decision explicit. My multi-tenant GPU orchestration on OpenShift AI post and the KubeCon Europe 2026 talk recap cover how this works in practice. GPU Sharing on Kubernetes: MIG vs MPS vs Time-Slicing compares the options.
3. Serve models behind a proper serving layer with SLOs. Pick a serving runtime (vLLM, Triton, NIM, or KServe on OpenShift AI), define latency SLOs such as time to first token, and benchmark against them before launch. See AI Model Serving on Kubernetes, Running LLMs on OpenShift AI, benchmarking vLLM with GuideLLM and autoscaling AI inference with HPA and KEDA.
4. Make it observable from day one. Collect GPU utilisation, token latency, throughput, error rates and cost per request in the same stack you use for everything else, and trace requests end to end. AI Observability on Kubernetes has the metrics and dashboards I start with.
5. Secure the path, not only the model. Use per-workload identity, least-privilege access to data and tools, private networking where the data requires it (the âSovereign by Designâ blueprint 2 above), and signed images. For agents, add approval gates, blast-radius limits and audit logs, as described in Guardrails for AI Agents in Production.
6. Turn governance into policy as code. Write down which models, registries, regions and data classes are allowed, and enforce those rules at admission time instead of relying on review meetings. Kyverno and policy-driven AI governance shows one way to do it on Kubernetes.
7. Put cost control in from the start. Tag every GPU workload with an owner and a cost centre, give teams showback, right-size and autoscale inference, and switch off idle capacity. FinOps for AI: Control GPU Costs has the playbook.
8. Give developers a golden path. Package all of the above as a template: a model service with CI, policies, dashboards and cost tags already wired in, available through an internal developer platform. That is where the âfaster, safer path from experimentation to productionâ comes from. See Internal Developer Platforms Compared.
If your team is stuck between a promising demo and a production rollout, this is the kind of work I do in AI integration consulting.