Tag:llm
All the articles with the tag "llm".
On AI Agents, Criminal Activity, And Who Is Actually Responsible
hrbrmstrOpenAI disclosed that GPT-5.6 Sol, running with cyber refusals disabled, broke out of its evaluation sandbox and compromised Hugging Face's production infrastructure. A zero-day in the proxy cache, lateral movement through OpenAI's research environment, stolen credentials chained with additional zero-days. The industry keeps framing cybersecurity as the headline AI risk for commercial reasons, but the real question isn't defense — it's liability. When someone configures and launches a model that commits crimes, do the humans who set it free bear responsibility? Computer fraud statutes were written with human actors in mind. This needs a test case.
Running Ornith Locally With OpenCode and Claude Code
hrbrmstrDeepReinforce's Ornith-1.0 models are solid agentic coders, but the upstream Ollama modelfile breaks tool calling out of the box. Here's the two-line fix, which variant fits your RAM, and how both the 35B and 9B performed on real honeypot triage work.
sdef2md: Turn any macOS app's scripting API into documentation and MCP tools
hrbrmstrA Go CLI that converts macOS .sdef scripting definitions into clean Markdown, paired with a skill that generates complete Go MCP servers from the generated reference — bridging any scriptable app into LLM agents.
Stop trusting LLM benchmarks
hrbrmstrEight major AI benchmarks can be gamed to near-perfect scores without solving tasks. Berkeley researchers show the scoring harnesses were never secure — and scores already inflated in the wild.