On 15 October 2024 I went to the Sustainable IT Netherlands Meetup at Accentureâs office in Amsterdam. There were two talks. Accenture spoke on âSustainable AI and Green codeâ, and Dr. Justus Bogner of the Vrije Universiteit Amsterdam presented âThe Challenge of Designing Green AI-Based Systems: A Software Architecture Perspectiveâ. Between them, the two talks went from âhow much carbon does a prompt cost?â down to âwhich design decisions make an ML system use less energy?â.
This is a throwback post, written from the photos I took of the slides.
Accenture: Sustainable AI and Green code
Where a GenAI modelâs footprint comes from
An early Accenture slide split the footprint of a GenAI model into three parts: training the model + inference + hardware. The training charts compared CO2 emissions, reported training time in days and power consumption for LLaMA, OPT, PaLM and BLOOM. The inference chart, âOperational Carbon Footprint of Large-Scale ML Tasksâ, put several of Metaâs production models next to published models such as T5, Meena, GShard, Switch Transformer and GPT-3.

Footprint = training + inference + hardware.
Five things to do when you build a GenAI solution
The practical slide had five points:
- Be discerning about when you use generative AI.
- Use smaller models when possible; fine-tuning brings good results.
- Re-use models and resources.
- Evaluate the energy sources of your cloud provider or data centre.
- Include the impact of AI activity in your carbon monitoring.

Five recommendations, starting with âbe discerning about when you use generative AIâ.
The next slide was demand shaping: âShape your computation to match the existing supply. If carbon intensity is low, increase the demand (do more in your applications). If carbon intensity is high, decrease demand (do less).â Itâs the idea behind carbon-aware scheduling, which I covered for Kubernetes in carbon-aware Kubernetes scheduling.

Demand shaping: do more when the grid is clean, do less when it isnât.
Measuring one prompt
Then came a measurement example. The tool was EcoLogits, an open-source library that estimates the environmental footprint of generative AI at inference time. The slide listed what it reports: energy consumption, global warming potential, abiotic depletion potential for elements, and primary energy. According to the slide, it covers OpenAI, Mistral AI, Google Gemini, LiteLLM âand moreâ. The code was a plain gpt-3.5-turbo chat completion asking for a joke. It then printed response.impacts.energy.value in kWh and response.impacts.gwp.value in kgCO2eq.

One joke from gpt-3.5-turbo, reported as a minâmax range of kWh and kgCO2eq.
The output was a range, not a single number: about 0.00009 to 0.0002 kWh, and about 0.00005 to 0.00012 kgCO2eq, for one short answer. Thatâs tiny per request. It isnât tiny once you multiply it by every request your product serves.
How green is AI-generated code?
The last part was about coding assistants. The slide cited a forecast that â70% of all developers will be using AI coding tools by 2027â, and argued that an assistantâs default coding behaviour is ânot greenâ. There were two levers: fine-tuning, so the model knows sustainable coding practices, and prompt engineering, giving it green-code prompts.

Default behaviour is ânot greenâ; fine-tuning and green prompts are the levers.
The customer story was âGreen Coding with GitHub Copilot for ACN Enterprise Searchâ. The team picked the top five AWS Lambda functions as hot spots, refactored them with âGreen Coding Promptsâ to GitHub Copilot, and measured execution times before and after. Execution times went from 1,400 to 608 ms for one function and from 956 to 267 ms for another, with smaller gains on the other three. The headline result on the slide was a 61.89% reduction in carbon emissions (MTCO2eq).
VU Amsterdam: green AI from a software architecture perspective

Dr. Justus Bognerâs title slide.
Justus Bogner introduced himself as an assistant professor in the Software and Sustainability (S2) group at VU Amsterdam. His starting point was that AI and software engineering have historically developed apart, and that âvery few people excel in both domainsâ. He then gave a reminder of what software architecture is for. It drives quality (not what the system does, but how), it lets teams work in parallel, and it is the bridge between requirements and implementation. Reusing architectural knowledge is how you raise quality without reinventing everything.
The energy argument followed. Training and using ML models is computationally demanding. The slides defined âGreen AIâ and âRed AIâ using the 2020 Communications of the ACM paper Green AI. Green AI yields new results while taking computational cost into account; Red AI buys accuracy with massive compute and disregards the cost. The âbad newsâ bullet was the honest one: itâs still difficult to apply the scattered research results in practice.
A catalogue of 30 green tactics
An architectural tactic is a high-level, reusable design decision aimed at one quality attribute, here energy efficiency. The study behind the talk has two stages. First, a literature synthesis of 51 solution papers from Verdecchia et al.âs 2023 Green AI review. Second, a focus group with three experts on architecture for ML-enabled systems, plus five more papers. The result is 30 green tactics for ML-enabled systems. The paper is A Synthesis of Green Architectural Tactics for ML-Enabled Systems by Heli JĂ€rvenpÀÀ, Patricia Lago, Justus Bogner, Grace Lewis, Henry Muccini and Ipek Ozkaya. It was presented at ICSE-SEIS 2024.

The full catalogue on one slide, from âapply sampling techniquesâ (T1) to âmonitor computing powerâ (T30).
The tactics are grouped by phase:
- Data-centric: apply sampling techniques (T1), remove redundant data, reduce the number of data features, use input quantisation, use data projection.
- Algorithm design: choose an energy-efficient algorithm (T6, with Naive Bayes, KNN, decision trees and linear models as examples), choose a lightweight alternative, decrease model complexity, use built-in library functions.
- Model optimisation: set energy consumption as a model constraint, enhance model sparsity, consider energy-aware pruning, transfer learning (T16) and knowledge distillation.
- Model training: use quantisation-aware training, use checkpoints, design for memory constraints.
- Deployment: consider federated learning, use computation partitioning, use energy-efficient hardware, use power capping, use energy-aware scheduling, minimise referencing to data.
- Management: use informed adaptation (T28, retraining on identified concept drift), retrain the model if needed, monitor computing power.
The slide titled âThe Best Green AI Tactic of Them Allâ was the simplest: only use AI for a problem when it actually makes sense. Citing Hultenâs Building Intelligent Systems, it described the ârightâ problems for AI: very large, open-ended, time-changing or intrinsically hard problem spaces. Imperfections must be acceptable, and AI must clearly beat the alternatives (heuristics, humans) once you include the cost of its mistakes. This matched Accentureâs first recommendation, âbe discerning about when you use generative AIâ, from a completely different angle.
The catalogue isnât only in the paper. Itâs on Zenodo and in a web app, the Archive of Awesome and Dark Tactics (AADT), run by VUâs S2 group, which has a category for green ML-enabled systems. The discussion slide warned that choosing tactics needs domain expertise, and that not every tactic applies everywhere, which is why some start with âconsiderâ and others with âuseâ. The final take-aways: software engineering for AI (SE4AI) is how you build high-quality AI-based systems, and architecture and design knowledge for them âis forming, but still in its infancyâ.
I saw the same green-tactics material again a month later, at the Mindstone AI Meetup Amsterdam in November 2024, which tells me the message was travelling beyond the sustainability crowd.
My take: measure energy like you measure cost
Both talks agreed that you canât improve what you donât measure. Both also showed how hard measuring is. Accentureâs Copilot story measured execution time and converted it to emissions. EcoLogits gives a modelled range for a request to an API whose hardware you never see. Neither is a meter reading. Both are still much better than nothing.
In platform work I treat energy the way I treat cloud spend:
- Start with a proxy you already have. CPU-seconds, GPU-hours and request counts are already in your metrics. Execution time, as in the Lambda example, is a fair first proxy for serverless and batch work.
- Attribute it. Showback by team and by feature is what changes behaviour, for kilowatt-hours as much as for euros. The same tagging and allocation work from FinOps for AI and GPU workloads gives you the energy view nearly for free.
- Measure on the cluster when you can. On Kubernetes, node and pod-level energy exporters plus carbon-intensity data let you act on demand shaping instead of just drawing it. I wrote about that in green Kubernetes with carbon-aware scheduling.
- Make the cheapest tactic the default. Smaller models, re-use, sampling, and âdonât use AI hereâ are architecture decisions. They belong in design reviews, not in a sustainability report after the fact.
The VU catalogue is useful here as a checklist. During an architecture review, go through the six groups and ask which tactics youâve considered and why you rejected the rest.