TIMESTAMP 2026-07-25T14:19:16,532671720+08:00 OLLAMA PROCESSES NAME ID SIZE PROCESSOR CONTEXT UNTIL shire-mini-fast:qwen3.5-4b 8c3d555cf1f7 3.1 GB 100% CPU 2048 4 minutes from now OLLAMA SERVICE LOGS — RECENT 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=patch_embed slice 0 (permute+F16->F32) tensor=v.patch_embd.weight bytes=3145728 duration_ms=6.274 total_ops=8 total_bytes=3198976 total_ms=6.453 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=patch_embed slice 1 (permute+F16->F32) tensor=v.patch_embd.weight.1 bytes=3145728 duration_ms=4.688 total_ops=9 total_bytes=6344704 total_ms=11.141 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=F16->F32 promote tensor=v.position_embd.weight bytes=9437184 duration_ms=14.140 total_ops=10 total_bytes=15781888 total_ms=25.281 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.0.attn_qkv.weight bytes=6291456 duration_ms=4.818 total_ops=11 total_bytes=22073344 total_ms=30.099 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.0.attn_qkv.bias bytes=12288 duration_ms=0.042 total_ops=12 total_bytes=22085632 total_ms=30.141 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.1.attn_qkv.weight bytes=6291456 duration_ms=4.915 total_ops=13 total_bytes=28377088 total_ms=35.055 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.1.attn_qkv.bias bytes=12288 duration_ms=0.039 total_ops=14 total_bytes=28389376 total_ms=35.094 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.2.attn_qkv.weight bytes=6291456 duration_ms=5.138 total_ops=15 total_bytes=34680832 total_ms=40.233 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.2.attn_qkv.bias bytes=12288 duration_ms=0.038 total_ops=16 total_bytes=34693120 total_ms=40.271 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.3.attn_qkv.weight bytes=6291456 duration_ms=5.053 total_ops=17 total_bytes=40984576 total_ms=45.324 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.3.attn_qkv.bias bytes=12288 duration_ms=0.043 total_ops=18 total_bytes=40996864 total_ms=45.366 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.4.attn_qkv.weight bytes=6291456 duration_ms=5.384 total_ops=19 total_bytes=47288320 total_ms=50.750 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.4.attn_qkv.bias bytes=12288 duration_ms=0.042 total_ops=20 total_bytes=47300608 total_ms=50.791 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.5.attn_qkv.weight bytes=6291456 duration_ms=4.757 total_ops=21 total_bytes=53592064 total_ms=55.548 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.5.attn_qkv.bias bytes=12288 duration_ms=0.040 total_ops=22 total_bytes=53604352 total_ms=55.588 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.6.attn_qkv.weight bytes=6291456 duration_ms=4.882 total_ops=23 total_bytes=59895808 total_ms=60.471 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.6.attn_qkv.bias bytes=12288 duration_ms=0.039 total_ops=24 total_bytes=59908096 total_ms=60.510 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.7.attn_qkv.weight bytes=6291456 duration_ms=4.845 total_ops=25 total_bytes=66199552 total_ms=65.355 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.7.attn_qkv.bias bytes=12288 duration_ms=0.047 total_ops=26 total_bytes=66211840 total_ms=65.402 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.8.attn_qkv.weight bytes=6291456 duration_ms=4.710 total_ops=27 total_bytes=72503296 total_ms=70.112 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.8.attn_qkv.bias bytes=12288 duration_ms=0.044 total_ops=28 total_bytes=72515584 total_ms=70.156 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.9.attn_qkv.weight bytes=6291456 duration_ms=4.826 total_ops=29 total_bytes=78807040 total_ms=74.983 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.9.attn_qkv.bias bytes=12288 duration_ms=0.046 total_ops=30 total_bytes=78819328 total_ms=75.029 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.10.attn_qkv.weight bytes=6291456 duration_ms=4.911 total_ops=31 total_bytes=85110784 total_ms=79.940 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.10.attn_qkv.bias bytes=12288 duration_ms=0.045 total_ops=32 total_bytes=85123072 total_ms=79.985 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.11.attn_qkv.weight bytes=6291456 duration_ms=4.912 total_ops=33 total_bytes=91414528 total_ms=84.897 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.11.attn_qkv.bias bytes=12288 duration_ms=0.046 total_ops=34 total_bytes=91426816 total_ms=84.943 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.12.attn_qkv.weight bytes=6291456 duration_ms=5.017 total_ops=35 total_bytes=97718272 total_ms=89.960 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.12.attn_qkv.bias bytes=12288 duration_ms=0.042 total_ops=36 total_bytes=97730560 total_ms=90.002 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.13.attn_qkv.weight bytes=6291456 duration_ms=4.703 total_ops=37 total_bytes=104022016 total_ms=94.705 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.13.attn_qkv.bias bytes=12288 duration_ms=0.046 total_ops=38 total_bytes=104034304 total_ms=94.751 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.14.attn_qkv.weight bytes=6291456 duration_ms=4.887 total_ops=39 total_bytes=110325760 total_ms=99.639 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.14.attn_qkv.bias bytes=12288 duration_ms=0.040 total_ops=40 total_bytes=110338048 total_ms=99.679 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.15.attn_qkv.weight bytes=6291456 duration_ms=4.833 total_ops=41 total_bytes=116629504 total_ms=104.512 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.15.attn_qkv.bias bytes=12288 duration_ms=0.047 total_ops=42 total_bytes=116641792 total_ms=104.560 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.16.attn_qkv.weight bytes=6291456 duration_ms=4.812 total_ops=43 total_bytes=122933248 total_ms=109.371 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.16.attn_qkv.bias bytes=12288 duration_ms=0.043 total_ops=44 total_bytes=122945536 total_ms=109.414 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.17.attn_qkv.weight bytes=6291456 duration_ms=4.726 total_ops=45 total_bytes=129236992 total_ms=114.140 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.17.attn_qkv.bias bytes=12288 duration_ms=0.043 total_ops=46 total_bytes=129249280 total_ms=114.183 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.18.attn_qkv.weight bytes=6291456 duration_ms=4.833 total_ops=47 total_bytes=135540736 total_ms=119.016 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.18.attn_qkv.bias bytes=12288 duration_ms=0.041 total_ops=48 total_bytes=135553024 total_ms=119.057 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.19.attn_qkv.weight bytes=6291456 duration_ms=4.845 total_ops=49 total_bytes=141844480 total_ms=123.902 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.19.attn_qkv.bias bytes=12288 duration_ms=0.040 total_ops=50 total_bytes=141856768 total_ms=123.943 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.20.attn_qkv.weight bytes=6291456 duration_ms=4.864 total_ops=51 total_bytes=148148224 total_ms=128.807 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.20.attn_qkv.bias bytes=12288 duration_ms=0.043 total_ops=52 total_bytes=148160512 total_ms=128.850 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.21.attn_qkv.weight bytes=6291456 duration_ms=5.240 total_ops=53 total_bytes=154451968 total_ms=134.089 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.21.attn_qkv.bias bytes=12288 duration_ms=0.041 total_ops=54 total_bytes=154464256 total_ms=134.130 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.22.attn_qkv.weight bytes=6291456 duration_ms=4.929 total_ops=55 total_bytes=160755712 total_ms=139.059 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.22.attn_qkv.bias bytes=12288 duration_ms=0.043 total_ops=56 total_bytes=160768000 total_ms=139.101 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.23.attn_qkv.weight bytes=6291456 duration_ms=5.169 total_ops=57 total_bytes=167059456 total_ms=144.271 2026-07-25T14:16:19+08:00 Work ollama[259962]: compat tensor transform: op=concat sources tensor=v.blk.23.attn_qkv.bias bytes=12288 duration_ms=0.040 total_ops=58 total_bytes=167071744 total_ms=144.311 2026-07-25T14:16:19+08:00 Work ollama[259962]: get_dummy_batch: warmup with image size = 1472 x 1472 2026-07-25T14:16:19+08:00 Work ollama[259962]: get_dummy_batch: warmup with image size = 1472 x 1472 2026-07-25T14:16:19+08:00 Work ollama[259962]: reserve_compute_meta: CPU compute buffer size = 223.30 MiB 2026-07-25T14:16:19+08:00 Work ollama[259962]: reserve_compute_meta: graph splits = 1, nodes = 736 2026-07-25T14:16:19+08:00 Work ollama[259962]: warmup: flash attention is enabled 2026-07-25T14:16:19+08:00 Work ollama[259962]: srv load_model: loaded multimodal model, '/usr/share/ollama/.ollama/models/blobs/sha256-81fb60c7daa80fc1123380b98970b320ae233409f0f71a72ed7b9b0d62f40490' 2026-07-25T14:16:23+08:00 Work ollama[259962]: cmn common_conte: the context does not support partial sequence removal 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv load_model: speculative decoding will use checkpoints 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv load_model: initializing, n_slots = 1, n_ctx_slot = 2048, kv_unified = 'false' 2026-07-25T14:16:23+08:00 Work ollama[259962]: spec common_specu: no implementations specified for speculative decoding 2026-07-25T14:16:23+08:00 Work ollama[259962]: slot load_model: id 0 | task -1 | new slot, n_ctx = 2048 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv load_model: prompt cache is enabled, size limit: 8192 MiB 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv load_model: use `--cache-ram 0` to disable the prompt cache 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv load_model: context checkpoints enabled, max = 32, min spacing = 8192 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv init: idle slots will be saved to prompt cache upon starting a new task 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv init: init: chat template, example_format: '<|im_start|>system 2026-07-25T14:16:23+08:00 Work ollama[259962]: You are a helpful assistant<|im_end|> 2026-07-25T14:16:23+08:00 Work ollama[259962]: <|im_start|>user 2026-07-25T14:16:23+08:00 Work ollama[259962]: Hello<|im_end|> 2026-07-25T14:16:23+08:00 Work ollama[259962]: <|im_start|>assistant 2026-07-25T14:16:23+08:00 Work ollama[259962]: Hi there<|im_end|> 2026-07-25T14:16:23+08:00 Work ollama[259962]: <|im_start|>user 2026-07-25T14:16:23+08:00 Work ollama[259962]: How are you?<|im_end|> 2026-07-25T14:16:23+08:00 Work ollama[259962]: <|im_start|>assistant 2026-07-25T14:16:23+08:00 Work ollama[259962]: ' 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv init: init: chat template, thinking = 0 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv llama_server: model loaded 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv llama_server: listening on http://127.0.0.1:40431 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv update_slots: all slots are idle 2026-07-25T14:16:23+08:00 Work ollama[259962]: time=2026-07-25T14:16:23.303+08:00 level=INFO source=llama_server.go:1274 msg="llama-server started in 11.82 seconds" 2026-07-25T14:16:23+08:00 Work ollama[259962]: time=2026-07-25T14:16:23.371+08:00 level=INFO source=images.go:373 msg="template selection" model=registry.ollama.ai/library/shire-mini-fast:qwen3.5-4b selected=renderer_parser renderer=qwen3.5 parser=qwen3.5 go_template=null chat_template="[tools thinking completion vision]" harmony=null renderer_parser="[completion vision tools thinking]" 2026-07-25T14:16:23+08:00 Work ollama[259962]: time=2026-07-25T14:16:23.372+08:00 level=INFO source=sched.go:739 msg="loaded runners" count=1 2026-07-25T14:16:23+08:00 Work ollama[259962]: time=2026-07-25T14:16:23.372+08:00 level=INFO source=llama_server.go:1207 msg="waiting for llama-server to start responding" 2026-07-25T14:16:23+08:00 Work ollama[259962]: time=2026-07-25T14:16:23.373+08:00 level=INFO source=llama_server.go:1274 msg="llama-server started in 11.89 seconds" 2026-07-25T14:16:23+08:00 Work ollama[259962]: slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv get_availabl: updating prompt cache 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 2048 tokens, 8589934592 est) 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv get_availabl: prompt cache update took 0.01 ms 2026-07-25T14:16:23+08:00 Work ollama[259962]: slot launch_slot_: id 0 | task -1 | sampler chain: logits -> penalties -> ?dry -> ?top-n-sigma -> top-k -> ?typical -> top-p -> ?min-p -> ?xtc -> temp-ext -> dist 2026-07-25T14:16:23+08:00 Work ollama[259962]: slot launch_slot_: id 0 | task -1 | sampler params: 2026-07-25T14:16:23+08:00 Work ollama[259962]: repeat_last_n = 64, repeat_penalty = 1.100, frequency_penalty = 0.000, presence_penalty = 1.500 2026-07-25T14:16:23+08:00 Work ollama[259962]: dry_multiplier = 0.000, dry_base = 1.750, dry_allowed_length = 2, dry_penalty_last_n = 2048 2026-07-25T14:16:23+08:00 Work ollama[259962]: top_k = 20, top_p = 0.950, min_p = 0.000, xtc_probability = 0.000, xtc_threshold = 0.100, typical_p = 1.000, top_n_sigma = -1.000, temp = 0.000 2026-07-25T14:16:23+08:00 Work ollama[259962]: mirostat = 0, mirostat_lr = 0.100, mirostat_ent = 5.000, adaptive_target = -1.000, adaptive_decay = 0.900 2026-07-25T14:16:23+08:00 Work ollama[259962]: slot launch_slot_: id 0 | task 0 | processing task, is_child = 0 2026-07-25T14:16:23+08:00 Work ollama[259962]: slot operator(): id 0 | task 0 | new prompt, n_ctx_slot = 2048, n_keep = 4, task.n_tokens = 33 2026-07-25T14:16:23+08:00 Work ollama[259962]: slot operator(): id 0 | task 0 | cached n_tokens = 0, memory_seq_rm [0, end) 2026-07-25T14:16:23+08:00 Work ollama[259962]: srv stream_sessi: conv_id= (empty=1) 2026-07-25T14:16:32+08:00 Work ollama[259962]: slot print_timing: id 0 | task 0 | prompt processing, n_tokens = 29, progress = 0.88, t = 9.06 s / 3.20 tokens per second 2026-07-25T14:16:32+08:00 Work ollama[259962]: slot operator(): id 0 | task 0 | cached n_tokens = 29, memory_seq_rm [29, end) 2026-07-25T14:16:32+08:00 Work ollama[259962]: slot init_sampler: id 0 | task 0 | init sampler, took 0.02 ms, tokens: text = 33, total = 33 2026-07-25T14:16:32+08:00 Work ollama[259962]: slot create_check: id 0 | task 0 | created context checkpoint 1 of 32 (pos_min = 28, pos_max = 28, n_tokens = 29, size = 50.251 MiB) 2026-07-25T14:16:54+08:00 Work ollama[259962]: slot print_timing: id 0 | task 0 | prompt eval time = 12661.90 ms / 33 tokens ( 383.69 ms per token, 2.61 tokens per second) 2026-07-25T14:16:54+08:00 Work ollama[259962]: slot print_timing: id 0 | task 0 | eval time = 18294.68 ms / 6 tokens ( 3049.11 ms per token, 0.33 tokens per second) 2026-07-25T14:16:54+08:00 Work ollama[259962]: slot print_timing: id 0 | task 0 | total time = 30956.58 ms / 39 tokens 2026-07-25T14:16:54+08:00 Work ollama[259962]: slot print_timing: id 0 | task 0 | graphs reused = 5 2026-07-25T14:16:54+08:00 Work ollama[259962]: slot release: id 0 | task 0 | stop processing: n_tokens = 38, truncated = 0 2026-07-25T14:16:54+08:00 Work ollama[259962]: srv update_slots: all slots are idle 2026-07-25T14:16:54+08:00 Work ollama[259962]: [GIN] 2026/07/25 - 14:16:54 | 200 | 43.450455054s | 127.0.0.1 | POST "/api/chat" 2026-07-25T14:16:57+08:00 Work ollama[259962]: slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = 341338187396 2026-07-25T14:16:57+08:00 Work ollama[259962]: srv get_availabl: updating prompt cache 2026-07-25T14:16:57+08:00 Work ollama[259962]: srv prompt_save: - saving prompt with length 38, total state size = 51.439 MiB (draft: 0.000 MiB) 2026-07-25T14:16:58+08:00 Work ollama[259962]: srv load: - looking for better prompt, base f_keep = 0.079, sim = 0.011 2026-07-25T14:16:58+08:00 Work ollama[259962]: srv update: - cache state: 1 prompts, 101.690 MiB (limits: 8192.000 MiB, 2048 tokens, 3061 est) 2026-07-25T14:16:58+08:00 Work ollama[259962]: srv update: - prompt 0x2563b140: 38 tokens, checkpoints: 1, 101.690 MiB 2026-07-25T14:16:58+08:00 Work ollama[259962]: srv get_availabl: prompt cache update took 56.07 ms 2026-07-25T14:16:58+08:00 Work ollama[259962]: slot launch_slot_: id 0 | task -1 | sampler chain: logits -> penalties -> ?dry -> ?top-n-sigma -> top-k -> ?typical -> top-p -> ?min-p -> ?xtc -> temp-ext -> dist 2026-07-25T14:16:58+08:00 Work ollama[259962]: slot launch_slot_: id 0 | task -1 | sampler params: 2026-07-25T14:16:58+08:00 Work ollama[259962]: repeat_last_n = 64, repeat_penalty = 1.100, frequency_penalty = 0.000, presence_penalty = 1.500 2026-07-25T14:16:58+08:00 Work ollama[259962]: dry_multiplier = 0.000, dry_base = 1.750, dry_allowed_length = 2, dry_penalty_last_n = 2048 2026-07-25T14:16:58+08:00 Work ollama[259962]: top_k = 20, top_p = 0.950, min_p = 0.000, xtc_probability = 0.000, xtc_threshold = 0.100, typical_p = 1.000, top_n_sigma = -1.000, temp = 0.000 2026-07-25T14:16:58+08:00 Work ollama[259962]: mirostat = 0, mirostat_lr = 0.100, mirostat_ent = 5.000, adaptive_target = -1.000, adaptive_decay = 0.900 2026-07-25T14:16:58+08:00 Work ollama[259962]: slot launch_slot_: id 0 | task 8 | processing task, is_child = 0 2026-07-25T14:16:58+08:00 Work ollama[259962]: slot operator(): id 0 | task 8 | new prompt, n_ctx_slot = 2048, n_keep = 4, task.n_tokens = 280 2026-07-25T14:16:58+08:00 Work ollama[259962]: slot operator(): id 0 | task 8 | checking checkpoint with [28, 28] against 3... 2026-07-25T14:16:58+08:00 Work ollama[259962]: slot operator(): id 0 | task 8 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 2026-07-25T14:16:58+08:00 Work ollama[259962]: slot operator(): id 0 | task 8 | erased invalidated context checkpoint (pos_min = 28, pos_max = 28, n_tokens = 29, n_swa = 0, pos_next = 0, size = 50.251 MiB) 2026-07-25T14:16:58+08:00 Work ollama[259962]: slot operator(): id 0 | task 8 | cached n_tokens = 0, memory_seq_rm [0, end) 2026-07-25T14:16:58+08:00 Work ollama[259962]: srv stream_sessi: conv_id= (empty=1) 2026-07-25T14:17:30+08:00 Work ollama[259962]: slot print_timing: id 0 | task 8 | prompt processing, n_tokens = 276, progress = 0.99, t = 32.67 s / 8.45 tokens per second 2026-07-25T14:17:30+08:00 Work ollama[259962]: slot operator(): id 0 | task 8 | cached n_tokens = 276, memory_seq_rm [276, end) 2026-07-25T14:17:30+08:00 Work ollama[259962]: slot init_sampler: id 0 | task 8 | init sampler, took 0.08 ms, tokens: text = 280, total = 280 2026-07-25T14:17:30+08:00 Work ollama[259962]: slot create_check: id 0 | task 8 | created context checkpoint 1 of 32 (pos_min = 275, pos_max = 275, n_tokens = 276, size = 50.251 MiB) 2026-07-25T14:18:23+08:00 Work ollama[259962]: [GIN] 2026/07/25 - 14:18:23 | 200 | 1m25s | 127.0.0.1 | POST "/api/chat" 2026-07-25T14:18:24+08:00 Work ollama[259962]: srv stop: cancel task, id_task = 8 2026-07-25T14:18:26+08:00 Work ollama[259962]: slot release: id 0 | task 8 | stop processing: n_tokens = 295, truncated = 0 2026-07-25T14:18:26+08:00 Work ollama[259962]: srv update_slots: all slots are idle 2026-07-25T14:18:27+08:00 Work ollama[259962]: slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.949 2026-07-25T14:18:27+08:00 Work ollama[259962]: slot launch_slot_: id 0 | task -1 | sampler chain: logits -> penalties -> ?dry -> ?top-n-sigma -> top-k -> ?typical -> top-p -> ?min-p -> ?xtc -> temp-ext -> dist 2026-07-25T14:18:27+08:00 Work ollama[259962]: slot launch_slot_: id 0 | task -1 | sampler params: 2026-07-25T14:18:27+08:00 Work ollama[259962]: repeat_last_n = 64, repeat_penalty = 1.100, frequency_penalty = 0.000, presence_penalty = 1.500 2026-07-25T14:18:27+08:00 Work ollama[259962]: dry_multiplier = 0.000, dry_base = 1.750, dry_allowed_length = 2, dry_penalty_last_n = 2048 2026-07-25T14:18:27+08:00 Work ollama[259962]: top_k = 20, top_p = 0.950, min_p = 0.000, xtc_probability = 0.000, xtc_threshold = 0.100, typical_p = 1.000, top_n_sigma = -1.000, temp = 0.000 2026-07-25T14:18:27+08:00 Work ollama[259962]: mirostat = 0, mirostat_lr = 0.100, mirostat_ent = 5.000, adaptive_target = -1.000, adaptive_decay = 0.900 2026-07-25T14:18:27+08:00 Work ollama[259962]: slot launch_slot_: id 0 | task 27 | processing task, is_child = 0 2026-07-25T14:18:27+08:00 Work ollama[259962]: slot operator(): id 0 | task 27 | new prompt, n_ctx_slot = 2048, n_keep = 4, task.n_tokens = 280 2026-07-25T14:18:27+08:00 Work ollama[259962]: slot operator(): id 0 | task 27 | checking checkpoint with [275, 275] against 279... 2026-07-25T14:18:27+08:00 Work ollama[259962]: slot operator(): id 0 | task 27 | restored context checkpoint (pos_min = 275, pos_max = 275, n_tokens = 276, n_past = 276, size = 50.251 MiB) 2026-07-25T14:18:27+08:00 Work ollama[259962]: slot operator(): id 0 | task 27 | cached n_tokens = 276, memory_seq_rm [276, end) 2026-07-25T14:18:27+08:00 Work ollama[259962]: srv stream_sessi: conv_id= (empty=1) 2026-07-25T14:18:27+08:00 Work ollama[259962]: slot init_sampler: id 0 | task 27 | init sampler, took 0.07 ms, tokens: text = 280, total = 280 2026-07-25T14:19:13+08:00 Work ollama[259962]: [GIN] 2026/07/25 - 14:19:13 | 200 | 46.27756962s | 127.0.0.1 | POST "/api/chat" 2026-07-25T14:19:14+08:00 Work ollama[259962]: srv stop: cancel task, id_task = 27 2026-07-25T14:19:16+08:00 Work ollama[259962]: [GIN] 2026/07/25 - 14:19:16 | 200 | 24.636µs | 127.0.0.1 | HEAD "/" 2026-07-25T14:19:16+08:00 Work ollama[259962]: [GIN] 2026/07/25 - 14:19:16 | 200 | 37.609µs | 127.0.0.1 | GET "/api/ps" 2026-07-25T14:19:17+08:00 Work ollama[259962]: slot release: id 0 | task 27 | stop processing: n_tokens = 295, truncated = 0 2026-07-25T14:19:17+08:00 Work ollama[259962]: srv update_slots: all slots are idle