ZEP-NAS: Enabling green-aware model design via zero-cost emission proxy in neural architecture search
Riccardo Cantini, Matteo Capalbo, Domenico Talia · Array · 2025
As the environmental impact of deep learning grows, sustainability must become a core design goal in model development. Yet, most existing approaches to sustainable Neural Architecture Search (NAS) address only the efficiency of the search process, while largely overlooking the substantial emissions generated by the resulting architectures during training and deployment. This gap is critical, as models found through NAS may remain computationally expensive and carbon-intensive despite efficient search. Integrating long-term sustainability directly into the NAS objective is therefore essential, but applying emission-based metrics remains challenging, as parallel evaluation of numerous candidate architectures across shared GPUs makes it difficult to attribute emissions to individual trials. To address these issues, we propose a novel NAS framework, ZEP-NAS ( Zero-cost Emission Proxy for Neural Architecture Search ), which leverages a Transformer-based in-context learning approach to estimate emissions in real time, enabling green-aware search without compromising parallelism. By optimizing a unified objective that balances performance and environmental impact, our method achieves substantial emission reductions across several datasets from the Computer Vision and NLP domains, attaining reductions of up to 43% with negligible drops in test accuracy, while maintaining near-linear scalability as the number of search jobs increases. These results highlight the importance of embedding sustainability into NAS objectives, ensuring that discovered models are both accurate and environmentally efficient. • Proposes ZEP-NAS, a green-aware NAS framework using zero-cost emission proxies. • Leverages a Transformer-based ICL approach for training-free emission estimation. • Cuts emissions by 40%+ with minimal accuracy loss, robust across CV and NLP tasks. • Shows near-linear scalability, achieving up to 21 × speed-up on 32 GPUs.