Problem
I don't have an effective means to write code using an LLM offline.
Goal
I need a system that would improve privacy with different projects. This is an investigation project which will be expanded to cover different LLM agents.
Motivation For Investigation
Different LLM privacy policies have provisions which may result in personal data being used to train models or share with third parties.
In four different privacy policies I reviewed, my concern is that there are weak provisions for intellectual property and private conversations whenever people use "AI".
As of August 14th, 2026:
- OpenAI https://openai.com/policies/row-privacy-policy/ states:
- they use user generated prompts to train their models but the data is "de-identified" in section 2 of their policy.
- personal data is shared with 3rd party vendors, business transfers, governments, etc in section 3.
- they may use your prompt data to train their models after 30 days. Even if you delete data off the website, they may keep the data you prompt for training.
- Deepseek https://cdn.deepseek.com/policies/en-US/deepseek-privacy-policy.html states:
- they may collect your text input, voice input, prompt, uploaded files, photos, feedback, chat history, or other content that you provide to our model
- personal data may be used to train their models.
- they are open with sharing personal data to 3rd parties.
- Copilot https://learn.microsoft.com/en-us/microsoft-365/copilot/microsoft-365-copilot-privacy states:
- user prompt data isn't used to train their LLM models.
- however includes a provision that there is a "filter for harmful content".
- their data privacy statement includes collecting data which "depends on the context of interactions" https://www.microsoft.com/en-ca/privacy/privacystatement.
- Anthropic https://www.anthropic.com/legal/privacy states:
- there is an opt-out feature for using personal data to train models.
- however, Anthropic collects and trains data from third parties.
- there are reports obligations to third parties with data transfers which may or may not include data transfers, and they not responsible for the decisions of third parties.
- much like other companies on this list, the data retention keeps personal data for "as long as reasonably necessary"
- similarly to OpenAI, they apply "aggregation and de-identification techniques" on personal data.
Investigation Notes
Current investigation:
- Initial Offline LLM (May 6th, 2026): My first interaction with an offline LLM was Gemma 4 EB, Qwen2.5 Coder 7B Instruct, and Mistral 7B instruct v0.3. Gemma 4 was the largest model of 7.5B parameters and takes 5.89GB of RAM. I was able to run these models to run on a computer with a Intel i5-6500, however, 1-2 tokens generated per second for each prompt.
- Mini Agent Development (May 10th, 2026): I created some small work flows to determine if an AI agent can code an application on its own by having a workflow Python script.
- Gemma 4 Workflows (May 12th, 2026): Gemma 4 was able to run small python scripts and smaller agent workflows.
- Increased Model Size (August 2nd, 2026): I was able to run Bansai 27B, Qwen3 Coder 30B A3B Instruct, and DeepSeek R1 Distill Qwen 32B on a better computer which generates 70 tokens per second on an RTX 5080/AMD processor. For coding applications, Bansai 27B and Qwen3 were able to perform better than other models.
- Hard Realization (August 9th, 2026): 30B models, while having a larger context window and better responses, tend to have an issue where modifications can't be made different code prompts. For example, I would ask one model to modify a datatype which would apply to several functions that I already created and the datatype would wouldn't change. A workflow is a little better but still can't create entire projects.
What I Learned
- Online AI usage policies and data retention raise concern about how intellectual property is handled.
- By comparison, an online coding agent like Claude is much better than offline coding agents for handling large tasks.
- A model like Gemma 4 is able to produce python scripts for smaller tasks and basic algorithms.
- If I use multiple smaller prompts for different tasks with an overall goal, I'm able to get better results than trying to create header code for an entire project.
- Larger models of roughly 30B parameters vs 7.5B parameters may be able to conceptualize ideas into code better in different programming languages, however, I found that I routinely have issues with Qwen3 which is routinely unable to make modifications to different algorithms.
Current Challenges
- Application Design: A lot of LLMs offline do not have consistent code output and are not able to strongly conceptualize applications. Designs still need to be worked out prior to creating different algorithms.
- Coding Agents: It appears that larger models can better report instructions for an agent to code programs. However, the JSON output prompt contains errors which a Python script needs to solve. LangGraph claims to have solved this issue online, however, this is still an online tool.
- Lack Of Defined Search Feature: I currently don't have a search tool built into my offline LLMs which means syntax for coding languages still needs to be searched up online.
Plan
- Software development still needs to be broken down into smaller well defined tasks. I need to create a simple workflow process which can apply for each project which utilizes the current abilities of offline LLM models. This will take about 2-3 months with my current work schedule.
- Workflows will need to be better defined. This may take 3-4 months on my task list to better understand an to create a smaller offline agent that can modify code correctly when problems show up. Automating errors and syntax in an IDE will need to be part of that stage.
- I'll need to use a search engine like SearXNG in the workflow process or download/vectorize coding standards for different programming languages that I work with.
This project isn't on my immediate task list and I'd like to get more information to better understand a good workflow so that Claude is no longer needed.