Skip to content

Manning – Build a DeepSeek Model (From Scratch) 2025

Updated August 10, 2026 22.3 MB
Manning – Build a DeepSeek Model (From Scratch) 2025

Download

About this item

Description

Build a DeepSeek Model (From Scratch) dissects the architecture of revolutionary DeepSeek models that have been able to perform on par with commercial giants at a fraction of the cost. The book’s main goal is to teach developers how to implement key innovations in the model, including mixed-expert (MoE) structures and multi-head hidden attention, so that they can build their own, efficient, and custom models.

Taking a step-by-step approach, the author starts with the basics of language models and then dives into the technical and engineering details that make DeepSeq so superior. Written specifically for machine learning engineers who want to push the boundaries of efficiency and accuracy in open source models, this book focuses on operational aspects and reducing training costs.

Book Features

  • Learning to implement the Mixture of Experts (MoE) architecture to increase efficiency.
  • Investigating the Multi-head Latent Attention (MLA) technique for memory optimization.
  • Training multi-token prediction (MTP) methods to speed up the inference process.
  • Guide to Model Distillation and Knowledge Transfer from Large to Small Models.
  • Implementing reinforcement learning loops to improve accuracy in programming tasks.
  • Presenting a four-step structure for reconstructing key features of the deep sea in the local environment.

Book specifications

  • Publisher: MANNING
  • Lecturer/Author: Raj Dandekar
  • Number of pages: 228
  • Number of chapters: 5
  • Format: PDF

Headlines

Build a DeepSeek Model (From Scratch)

Pictures

User Guide

Extract the file and run it with the appropriate software.

Download link

Download file – 22.3 MB

File(s) password: www.downloadly.ir

File size

22.3 MB

What is included

  • Learning to implement the Mixture of Experts (MoE) architecture to increase efficiency.
  • Investigating the Multi-head Latent Attention (MLA) technique for memory optimization.
  • Training multi-token prediction (MTP) methods to speed up the inference process.
  • Guide to Model Distillation and Knowledge Transfer from Large to Small Models.
  • Implementing reinforcement learning loops to improve accuracy in programming tasks.
  • Presenting a four-step structure for reconstructing key features of the deep sea in the local environment.