by langgenius · v0.0.17
Lemonade lets you run local models with NPU and GPU acceleration.
This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.
Available inside your emploidai workspace after installation.
Available inside your emploidai workspace after installation.
Lemonade is a client inference framework (Windows, Linux) designed for seamless deployment of large language models (LLMs) with NPU and GPU acceleration. It supports models like Qwen, Llama, DeepSeek, and more, optimized for different hardware configurations.
Lemonade enables local execution of LLMs, providing enhanced data privacy and security by keeping your data on your own machine while leveraging hardware acceleration for improved performance.
Dify integrates with Lemonade Server to provide LLM (including vision and structured output), text embedding, reranking, speech-to-text (Whisper), and text-to-speech (Kokoro) capabilities for models deployed locally.
Visit the Lemonade's Website to download the Lemonade Server client for your system.
To start Lemonade server, simply run:
lemonade-server serve
Note: If you installer from source use
lemonade-server-devinstead
Once started, Lemonade will be accessible at http://localhost:13305.
Go to the Dify marketplace, search for "Lemonade", and click to install the official plugin.
If this is your first time running Dify, you can find out how to get started here.
Go to Settings > Model Providers > Lemonade > Add a Model.
Then, fill in the following configuration:
Basic Configuration:
llm, text-embedding, rerank, speech2text, or tts based on your use casehttp://127.0.0.1:13305http://192.168.1.100:13305 or http://host.docker.internal:13305LLM options:
Speech-to-Text options (Whisper recipes):
en, zh).Text-to-Speech options (Kokoro recipes):
mp3 or wav).Sample Configuration:
Model Name: Qwen3-8B-GGUF
Model Type: llm
Authorization Name: (leave blank)
Context Size: 4096
Agent Thought: Support
Vision Support: Not Support
You can now use Lemonade with your favorite Dify workflow!
You can manage which models are installed on Lemonade using the Model Management GUI.
http://localhost:13305If you are using an AMD RyzenAI 300 series processor, you are able to use NPU and Hybrid (NPU+iGPU) acceleration. This include models like Llama-3.1-8B-Instruct-Hybrid and many others.
For a complete list of supported models, visit Lemonade Server Models.
Lemonade contains a series of advanced options, including server-level context size configurations, and Llama.cpp ROCm support. For aditional details, please check Lemonade Server Documentation and the Lemonade Server GitHub Repository.