Course Introduction


Model servers are critical to any AI solution today, but many applications don’t need the power, or cost, of cloud-based inference. A private model server tailored to your specific problem can work just as well, if not better, while being less expensive to run and deploy. Better still, you can build an application that interacts with models directly, with no server at all.

Projects like Yzma and Kronk make this a reality. Learn to write software that interacts with open-source models directly, and build your own model server when you need one. Use the Kronk SDK for hardware-accelerated local inference with llama.cpp integrated directly into your Go applications via Yzma, with no CGO required.

This course is part of the Ultimate Go Track. Not sold separately.

Note: All of our bundles are for a one-year subscription.

At the end of the subscription period, your membership does not automatically renew.

Requirements:

Students with the following background will get the most out of the course:

  • Several months of experience coding in Go
  • A working Go environment on the device you will use for the course

You will need one of the following to run the models:

  • A Mac with an M1 series chip or newer and at least 16GB of RAM (32GB+ preferred)
  • A Linux or Windows laptop with a dedicated GPU with at least 8GB of VRAM, not system RAM (16GB preferred)
  • Access to a cloud-based instance with a dedicated GPU with at least 8GB of VRAM (16GB preferred)

Topics Covered


Recorded from a live session, this course covers the following topics:

  • Using llama.cpp libraries directly in Go applications via the Yzma project
  • Building robust Go applications that interact with open-source models from Hugging Face via Kronk
  • Hardware-accelerated local inference with the Kronk SDK, with no CGO required
  • Building a model server with an OpenAI-compatible chat completions API for tailored needs

Skills You’ll Gain


  • Local model inference
  • llama.cpp in Go
  • Yzma
  • Kronk SDK
  • Hugging Face open-source models
  • Model servers
  • OpenAI-compatible APIs