vllm
FeaturedDeploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads.
Install
Quality Score: 93/100
Skill Content
Details
- Author
- Prism-Shadow
- Repository
- Prism-Shadow/penguin-harness
- Created
- 1 months ago
- Last Updated
- today
- Language
- TypeScript
- License
- Apache-2.0
Integrates with
Similar Skills
Semantically similar based on skill content — not just same category
ollama
Deploy and serve local models with Ollama — pull and run them, then expose the OpenAI-compatible endpoint to apps and agents.
local-llm-agent
Use the high-end LOCAL LLMs on this machine as a real agent/coding substrate (not just Claude). Dual RTX 5090 (64GB VRAM) + ollama already installed and serving. Best agentic-coding local model: qwen3-coder:30b. Covers how to pull, run, call (CLI / HTTP / OpenAI-compatible / function-calling), route through the AIOS provider harness, and use as a heterogeneous arm in absorption-probe. Per feedback_use_all_substrates_not_own_head — don't solve from one model.
ai-local-model-ops
Runs local and self-hosted LLM workflows with Ollama, LM Studio, MLX, Open WebUI, llamafile, and adapters. Use when operating private model stacks.