by wanghualoong · v0.1.4
Connect Dify to a self-hosted New API gateway (QuantumNous/new-api), an OpenAI/Claude/Gemini-compatible model relay. Fixes Qwen multimodal input, embedding dimensions and speech-to-text compatibility.
This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.
Available inside your emploidai workspace after installation.
Available inside your emploidai workspace after installation.
Connect Dify to a self-hosted New API gateway — an OpenAI/Claude/Gemini-compatible AI model relay and asset-management system. New API normalizes many upstream providers into OpenAI-compatible endpoints, so a single connection exposes chat, embedding, rerank, speech-to-text and text-to-speech models from whichever channels your gateway is configured with.
This plugin is purpose-built for New API and fixes several compatibility gaps that the generic OpenAI-API-compatible provider has with models routed through the gateway — especially Qwen / Dashscope.
| Type | Endpoint used | Notes |
|---|---|---|
| LLM | /v1/chat/completions (or /v1/completions) | Covers OpenAI, Claude and Gemini chat models, which New API exposes in OpenAI format. Supports vision, audio, video and document input, tool/function calling, reasoning/thinking mode, structured output and an optional web-search toggle. |
| Text Embedding | /v1/embeddings | Configurable dimensions and per-request batch size for Dashscope text-embedding-v3/v4 and similar. |
| Rerank | /v1/rerank | Jina / Cohere / Xinference style. |
| Speech2Text | /v1/audio/transcriptions | language and prompt are optional, so gpt-4o-transcribe / qwen-asr work. |
| TTS | /v1/audio/speech | Configurable voices and audio format. |
Image, video and music generation are not Dify model-provider types and are out of scope for this plugin.
input_audio part ({data, format}) instead of being wrapped in image_url.video_url (Qwen-VL) instead of image_url.openai_file (OpenAI file_data), qwen_fileid (upload via /v1/files and reference fileid://, for Qwen-Long / Dashscope), or image_fallback (render as an image).dimensions parameter and small batch caps, with an automatic single-input fallback when a batch is rejected.language/prompt.Models are added individually (customizable models), so each entry maps to a real model on your gateway and stays fully editable afterwards.
/v1, e.g. https://your-newapi-domain/v1 (default http://localhost:3000/v1).Set Thinking Mode Support to Only Non-Thinking Mode to suppress reasoning. The plugin then sends the common vendor "disable thinking" hints (enable_thinking, thinking: {type: disabled}, chat_template_kwargs) so models that support a toggle (e.g. Qwen3, GLM, Doubao) stop generating reasoning at the source, and it also strips any reasoning that still leaks into the answer (<think>...</think> or a lone </think>). Note: some dedicated reasoning models (e.g. DeepSeek-R1) always reason and cannot be disabled — their reasoning is hidden from the answer but still generated.
When credential validation fails because a model name is not found, the error message lists the models reported by your gateway's /v1/models endpoint. You can also query it directly:
curl https://your-newapi-domain/v1/models -H "Authorization: Bearer <your-token>"
Note: Dify's plugin framework does not support fetching a remote model list into the configuration form, so models are added individually. To bulk-create entries for every model your gateway serves, read its
/api/pricinglisting and create the models via the Dify Console API.
For embedding models with Vision Support enabled (e.g. Qwen3-VL-Embedding), inputs may carry images/video in addition to text. Each input can be a JSON object {"text": "...", "image": "<url|data-uri>"} (or "video"), a bare image/video URL, data: URI, or markdown image ; plain strings are sent as text. When vision is enabled every input is sent to /v1/embeddings as an object (the gateway rejects lists that mix plain strings and objects). Note that Dify's standard knowledge-base RAG only embeds text — image/video embedding applies when a caller supplies such inputs explicitly.
This plugin's source is provided under the same license as the Dify plugin ecosystem. "New API" and its logo belong to the New API project (QuantumNous/new-api); they are used here only to identify the gateway this plugin connects to.