Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I just downloaded llama. ran this llama-cli -hf ggml-org/gemma-3-1b-it-GGUF

I am getting [ Prompt: 91.3 t/s | Generation: 171.8 t/s ]

This is on a GPU (RTX 4060)

Is this decent?



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: