SRE for AI Systems Pravin Nair
- Autor:
- Pravin Nair
- Wydawnictwo:
- BPB Publications
- Ocena:
- Stron:
- 308
- Dostępne formaty:
-
ePubMobi
Opis
książki
:
SRE for AI Systems
As AI and ML rapidly power modern digital services from recommendation engines to generative models, moving these workloads into production exposes critical gaps in traditional operations. Site reliability engineering (SRE) applies software engineering principles to infrastructure, making SRE uniquely positioned to own the reliability, observability, and resilience of complex AI-driven environments.
This book bridges the gap between traditional SRE practices and the innovative strategies required to manage AI-driven infrastructures such as LLM models effectively, offering readers a focused guide for this transformative era. Each chapter walks the reader through core concepts such as versioning of models and data, pipeline monitoring, and security testing for AI APIs, while providing concrete SRE patterns for uptime, rollback, and incident management in AI-driven environments. You will learn how to design service-level objectives that reflect AI-specific quality metrics, implement robust monitoring and alerting for model drift and data quality, secure AI-backed APIs against prompt injection and other attacks, and coordinate releases across code, data, and models in a complex environment.
By the end of this book, as an SRE engineer, you will be equipped to design, operate, and scale AI systems with the same rigor that you applied to traditional infrastructure. You will gain skills in building observable AI pipelines, managing versioned data and models at scale, securing AI APIs, and applying SRE principles such as error budgets, incident response, and automation to complex ML and generative AI workloads. What you will learn
Applying SRE principles to AI and ML workloads.
Monitoring AI pipelines for model quality, data drift, and output correctness.
Defining SLIs and SLOs that reflect business metrics.
Building resilient deployment and rollback strategies for LLM models.
Equipping SRE teams to own AI pipelines and ML lifecycle.
Building and scaling SRE teams to manage AI systems. Who this book is for
This book is for site reliability engineers, ML engineers, software developers, and SRE managers supporting AI workloads. Readers should possess basic cloud infrastructure knowledge, a foundational understanding of software development, and familiarity with general machine learning principles. Table of Contents
1. Introduction to SRE and AI Systems
2. Reliability Challenges in AI Workloads
3. Reliability Failures in AI Systems
4. Monitoring and Observability for AI Systems
5. AI-enhanced Automation in SRE
6. Resilient Architecture in Cloud
7. Incident Management and Root Cause Analysis with AI
8. SLOs and Error Budgets for AI Systems
9. CI/CD and Testing Using AI and SRE Principles
10. Building and Scaling AI SRE Teams
11. Cultural Shifts and Collaboration in AI SRE
12. Future Trends and Ethical AI in SRE
Dzięki opcji "Druk na żądanie" do sprzedaży wracają tytuły Grupy Helion, które cieszyły sie dużym zainteresowaniem, a których nakład został wyprzedany.
Dla naszych Czytelników wydrukowaliśmy dodatkową pulę egzemplarzy w technice druku cyfrowego.
Co powinieneś wiedzieć o usłudze "Druk na żądanie":
- usługa obejmuje tylko widoczną poniżej listę tytułów, którą na bieżąco aktualizujemy;
- cena książki może być wyższa od początkowej ceny detalicznej, co jest spowodowane kosztami druku cyfrowego (wyższymi niż koszty tradycyjnego druku offsetowego). Obowiązująca cena jest zawsze podawana na stronie WWW książki;
- zawartość książki wraz z dodatkami (płyta CD, DVD) odpowiada jej pierwotnemu wydaniu i jest w pełni komplementarna;
- usługa nie obejmuje książek w kolorze.
Masz pytanie o konkretny tytuł? Napisz do nas: [email protected]
Książka drukowana

Oceny i opinie klientów: SRE for AI Systems Pravin Nair
(0)