CodeBench: mallikohtainen golden example (profiles.json → golden kenttä)
qwen3-coder:30b → todo.md (annotaatiot) qwen3:8b → todo-readme.md (GitHub README -muoto, tutuin koulutusdata) Golden example ladataan dynaamisesti per malli pipelinen sisällä.
This commit is contained in:
69
kipina-codebench/results/2026-04-14T10-59.json
Normal file
69
kipina-codebench/results/2026-04-14T10-59.json
Normal file
@@ -0,0 +1,69 @@
|
||||
[
|
||||
{
|
||||
"model": "qwen3:8b",
|
||||
"scenario": "blog",
|
||||
"reqOk": true,
|
||||
"specOk": true,
|
||||
"specEntities": 2,
|
||||
"validationIssues": 0,
|
||||
"fixRounds": 1,
|
||||
"testsTotal": 11,
|
||||
"testsPassed": 11,
|
||||
"testsFailed": 0,
|
||||
"totalDurationMs": 64124,
|
||||
"totalTokens": 5689,
|
||||
"avgTokPerSec": 98.61378134916481,
|
||||
"promptChars": 12098,
|
||||
"promptTokensEst": 3025,
|
||||
"score": 90,
|
||||
"stars": "★★★★★",
|
||||
"error": null,
|
||||
"profile": "small",
|
||||
"promptName": "code-small",
|
||||
"round": 1
|
||||
},
|
||||
{
|
||||
"model": "qwen3:8b",
|
||||
"scenario": "blog",
|
||||
"reqOk": true,
|
||||
"specOk": true,
|
||||
"specEntities": 2,
|
||||
"validationIssues": 0,
|
||||
"fixRounds": 3,
|
||||
"testsTotal": 0,
|
||||
"testsPassed": 0,
|
||||
"testsFailed": 0,
|
||||
"totalDurationMs": 126014,
|
||||
"totalTokens": 11162,
|
||||
"avgTokPerSec": 97.09858655726343,
|
||||
"promptChars": 12101,
|
||||
"promptTokensEst": 3025,
|
||||
"score": 0,
|
||||
"stars": "☆☆☆☆☆",
|
||||
"error": "Testit kaatuivat",
|
||||
"profile": "small",
|
||||
"promptName": "code-small",
|
||||
"round": 2
|
||||
},
|
||||
{
|
||||
"model": "qwen3:8b",
|
||||
"scenario": "blog",
|
||||
"reqOk": true,
|
||||
"specOk": false,
|
||||
"specEntities": 0,
|
||||
"validationIssues": 0,
|
||||
"fixRounds": 0,
|
||||
"testsTotal": 0,
|
||||
"testsPassed": 0,
|
||||
"testsFailed": 0,
|
||||
"totalDurationMs": 0,
|
||||
"totalTokens": 0,
|
||||
"avgTokPerSec": 0,
|
||||
"promptChars": 0,
|
||||
"promptTokensEst": 0,
|
||||
"score": 0,
|
||||
"stars": "",
|
||||
"error": "JSON-speksi epäonnistui",
|
||||
"round": 3
|
||||
}
|
||||
]
|
||||
Reference in New Issue
Block a user