vllm · emploidai Marketplace
emploidai Marketplace
Add-onsAppletsPlugins
Search tools, teams, and capabilitiesPublish
MarketplacePluginsvllm
Plugin
Limited listing

vllm

by yangyaofei · v0.2.3

vllm provider for extra_body support https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html#id5

95k installsUpdated May 13, 2026
Publisher information is incomplete

This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.

Capabilities

Models

Available inside your emploidai workspace after installation.

Data sources

Available inside your emploidai workspace after installation.

Category

model

Version

0.2.3yangyaofei

Requirements

Maximum memory 256MB

Pricing

Not disclosed by publisher

Security & access

Review before installing

CompatibleRequires emploidai 1.0.0+

Permissions

  • Uses model capability

Dependencies

No additional dependencies

emploidai Marketplace

Discover capabilities. Review access. Install inside your workspace.

DocumentationSecuritySupportPrivacyTerms

vLLM Dify Provider

Dify custom model provider for vLLM's OpenAI-Compatible Server, supporting extra parameters and thinking mode features.

Based on the official Dify OpenAI-API-compatible plugin, extended for vLLM OpenAI-Compatible Server.

Latest: v0.2.3

Fixes

  • #34: Fixed thinking mode markup tags — <think/>/</think/> → <think>/</think>, aligned 1:1 with dify-official-plugins openai_api_compatible
  • #31: Fixed extra_body parameter delivery — JSON contents now merged into top-level request body instead of nested under "extra_body" key

Features (since v0.2.0)

  • extra_body: Pass any vLLM extra parameters directly as JSON
  • Thinking Mode: enable_thinking toggle with compatibility_mode (strict/extended)
  • Reasoning Effort: Support reasoning_effort (none/low/medium/high), natively supported by vLLM
  • Compatibility Mode: Extended mode injects chat_template_kwargs, thinking, enable_thinking at top level
  • Structured Output: Support response_format, json_schema, reasoning_format
  • Thinking Content Filter: Auto-filter <think>...</think> when thinking is disabled
  • Thinking Content Cleanup: Strip thinking content from history before requests
  • vLLM Reasoning Field: Priority read reasoning (vLLM >= 0.17.1), fallback to reasoning_content

v0.2.0 Baseline

Breaking Change: Removed all legacy parameters, use extra_body for all extra parameter needs.

Usage

Add model

Same as OpenAI-API-compatible, select "Vllm" provider:

Configure extra_body

Pass extra parameters via the extra_body JSON text field:

Example:

{"chat_template_kwargs": {"enable_thinking": true}}

Repo

https://github.com/yangyaofei/dify-vllm-provider