Hi!
I'm trying to switch from Locally AI and found out Atomic Chat, which seemed to be exactly what i was searching for, clean, fast and responsive UI, and open source to be sure no data is tracked.
But while benchmarking it against Locally AI (from the LM Studio team) i noticed that on the same model (Qwen3.5 4B) it is quite slower 20% up to 40% on extremely long prompt and long responses.
That's the question, doesn't Atomic Chat uses MLX models on iPhone/iPad like Locally? It seems to use GGUF, i noticed there is the "import GGUF" but i thought that the "dafault" provided models were all MLX which should give more performance on iOS.
Also for the suggestion would be awesome to have more details for the "default" models, like the quantization, era those all Q4_K_M? like a modal or just some more info under the model name (and as stated before if it is a MLX or GGUF).
Thanks a lot for your amazing work.
Hi!
I'm trying to switch from Locally AI and found out Atomic Chat, which seemed to be exactly what i was searching for, clean, fast and responsive UI, and open source to be sure no data is tracked.
But while benchmarking it against Locally AI (from the LM Studio team) i noticed that on the same model (Qwen3.5 4B) it is quite slower 20% up to 40% on extremely long prompt and long responses.
That's the question, doesn't Atomic Chat uses MLX models on iPhone/iPad like Locally? It seems to use GGUF, i noticed there is the "import GGUF" but i thought that the "dafault" provided models were all MLX which should give more performance on iOS.
Also for the suggestion would be awesome to have more details for the "default" models, like the quantization, era those all Q4_K_M? like a modal or just some more info under the model name (and as stated before if it is a MLX or GGUF).
Thanks a lot for your amazing work.