ZooWork
ZooWork Market · Skill

llama-cpp

Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. Use for edge deployment, M1/M2/M3 Macs, AMD/Intel GPUs, or when CUDA is unavailable. Supports GGUF quantization (1.5-8 bit) for reduced memory and 4-10× speedup vs PyTorch on CPU.

◇
47thstreet
47thstreet-clawdagent-llama-cpp · v1.0.0
分类ai-llms
安装次数8
更新时间2026-08-05T00:04:46.624Z
校验状态待验证
Inference ServingLlama.cppCPU InferenceApple SiliconEdge DeploymentGGUFQuantizationNon-NVIDIAAMD GPUsIntel GPUsEmbedded