Description
Production LLM Deployment: vLLM, FastAPI, Modal and AI Chatbot. This comprehensive course teaches participants how to efficiently and scalablely deploy large language models to support thousands of requests, leveraging cutting-edge technologies such as vLLM, FastAPI and the Modal platform. Combining theoretical foundations with intensive hands-on exercises, this course provides the skills needed to design and launch robust, interactive inference services. Participants learn how to deploy models capable of processing thousands of parallel requests using vLLM, design a modular pipeline for model download and inference, and create a memory-capable chatbot that interacts with built-in endpoints. Additionally, they develop powerful, OpenAI-compatible APIs using FastAPI and vLLM, and deploy them in containerized environments to facilitate integration with external applications. Optimal resource management through concurrency and synchronization techniques for high model availability, optimizing GPU usage for parallel requests, and designing secure and scalable services with advanced authentication mechanisms and token-based access control are also covered. This course provides an in-depth introduction to the Modal platform, which, with its infrastructure-as-code approach, simplifies the deployment and scaling of models with just a few Python decorators. Topics include setting up an environment, running local and remote scripts, deploying a model with a single command, managing the application lifecycle, converting functions to web services with FastAPI, using lifecycle hooks to manage resources, and fully automating the deployment workflow. Ultimately, this course will enable participants to deploy AI models for business applications, personal projects, or customer engagements in a professional and efficient manner.
What you will learn
- They will learn volume mapping to efficiently manage model storage, reduce redundant data retrieval, optimize weight storage, and speed up access using a local storage strategy.
- They learn to deploy AI models with vLLM, manage thousands of requests, and design modular architectures for efficient model download and inference.
- They will learn to create a conversational AI chatbot using Python, integrating the OpenAI API for real-time, seamless chats with established language models.
- Learn to use FastAPI and vLLM to build efficient and OpenAI-compatible APIs; deploy REST API endpoints in containers for seamless AI model interactions with external applications.
- They learn to use concurrency and synchronization to manage models, ensure high availability, and optimize GPU usage to efficiently handle many parallel inference requests.
- They will learn to design scalable systems with efficient scaling through Local Model Weights and Storage; and secure applications using Advanced Authentication and Token-based Access Control.
- They learn to execute GPU or CPU intensive functions in an application running locally on a powerful remote infrastructure (Modal Remote Infrastructure).
- They will learn to deploy AI models with a single command to run on a remote infrastructure defined in the application code.
- Implementing Web APIs: Learn to convert Python functions into Web Services using FastAPI in Modal, effectively integrating with multi-language applications.
This course is suitable for people who:
- This course is designed for software developers and IT professionals looking to improve their skills in deploying and scaling machine learning models in a cloud environment.
- People who want to move beyond traditional infrastructure challenges like manual scaling and complex server setups and are interested in using serverless architecture for simplified operations.
- Participants who appreciate a hands-on approach to learning and focus on implementing real-world solutions including API integration, container management, and cost-effective deployment strategies.
- People who want to deepen their understanding of cloud-based technologies, especially when it comes to optimizing machine learning workflows using platforms like Modal.
Course details
- Publisher: Udemy
- Instructor: Petar Petkanov
- Training level: Beginner to advanced
- Training duration: 5 hours and 28 minutes
- Number of lessons: 20
Course headings
Prerequisites for the Production LLM Deployment: vLLM FastAPI Modal and AI Chatbot course
- Basic Python Skills: Familiarity with Python programming, as the course involves scripting and using Python-based tools.
- Understanding of Machine Learning Concepts: A foundational grasp of machine learning principles and workflows will help in the application of deployment strategies.
- Experience with Command Line Interfaces: Competence in using command line tools for installing packages and running scripts is beneficial.
- Access to a Computer with Internet: A reliable computer setup with Internet access is necessary to follow along with the cloud-based exercises and deployments.
Course images
Sample course video
Installation Guide
After Extract, view with your favorite player.
Subtitles: English
Quality: 720p
Download link
Downloadly
Rapidgator link
File(s) password: www.downloadly.ir
File size
4.4 GB

