09-03-2026, 11:43 PM
Incidentally, one of my coworkers runs Qwen ("open-weight") on his home PC. I thought you needed one of those $16,000 cards to even have a start point but some of the Chinese models will run locally (that is, they are not running on a Chinese server). Personally I would not trust a Chinese made AI model no matter what anyone says but he seemed unconcerned about it, and you could always create a defensive setup (like make it run in a docker container with no network, or a tightly screened network).
Anyway, he says it can do about 1.5 tokens per second (which is dead slow). It basically works about as fast as a person could type. But he says it actually is pretty smart. And free. He gives it a coding task and just lets it go. Takes hours to do something Claude would do in 10 minutes, but it gets there, all running locally. The fact that this works at all is kinda wild. I thought we were a long way away from having AI that didn't run on a billion dollar server farm but maybe we're closer than I realize.
Anyway, he says it can do about 1.5 tokens per second (which is dead slow). It basically works about as fast as a person could type. But he says it actually is pretty smart. And free. He gives it a coding task and just lets it go. Takes hours to do something Claude would do in 10 minutes, but it gets there, all running locally. The fact that this works at all is kinda wild. I thought we were a long way away from having AI that didn't run on a billion dollar server farm but maybe we're closer than I realize.
