52 lines
2.1 KiB
YAML
52 lines
2.1 KiB
YAML
eval_type: inference-check
|
|
model: Nanthasit/sakthai-plus-1.5b
|
|
timestamp: 2026-07-31T00:29:43Z
|
|
results:
|
|
- method: curl POST to api-inference.huggingface.co
|
|
status: dns_unreachable
|
|
detail: "api-inference.huggingface.co does not resolve in sandbox DNS (gaierror -5)"
|
|
http_status: null
|
|
response_time_sec: null
|
|
output: null
|
|
|
|
- method: InferenceClient with provider='auto'
|
|
status: no_provider_mapping
|
|
detail: "Model has empty inference_provider_mapping; StopIteration in provider selection"
|
|
http_status: null
|
|
response_time_sec: 0.134
|
|
output: null
|
|
|
|
- method: InferenceClient with provider='hf-inference'
|
|
status: model_not_supported
|
|
detail: "BadRequestError: Model not supported by provider hf-inference"
|
|
http_status: 400
|
|
response_time_sec: 0.265
|
|
output: '{"error":"Model not supported by provider hf-inference"}'
|
|
|
|
- method: router.huggingface.co/v1/chat/completions
|
|
status: model_not_supported
|
|
detail: "Model not supported by any enabled provider"
|
|
http_status: 400
|
|
response_time_sec: 0.148
|
|
output: '{"error":{"message":"The requested model is not supported by any provider you have enabled.","code":"model_not_supported"}}'
|
|
|
|
- method: local transformers inference
|
|
status: oom
|
|
detail: "OOM (exit 137) — sandbox has 7.8GB RAM, 712MB free; 1.5B model requires ~3GB (fp16) or ~6GB (fp32)"
|
|
http_status: null
|
|
response_time_sec: null
|
|
output: null
|
|
|
|
summary:
|
|
accessible: false
|
|
root_cause: |
|
|
The old inference API (api-inference.huggingface.co) is fully deprecated and has no DNS records.
|
|
The new Inference Providers router rejects the model because no provider has it in their catalog.
|
|
Local inference impossible due to memory constraints (712MB free).
|
|
recommendation: |
|
|
To get this model serving inference, either:
|
|
a) Enable serverless inference for the model on HF Hub (Settings → Inference), which makes
|
|
hf-inference provider load it on-demand.
|
|
b) Deploy a dedicated Inference Endpoint ($$ — not compatible with Zero-Cost First principle).
|
|
c) Convert to GGUF and run via llama.cpp on a machine with ≥4GB RAM.
|