CPU-Only Self Enhancing Authoring Copilot Design-Based Markov Decision Processes Orchestration and Qwen 3 Local Large Language Model

We introduce a novel, privacy-preserving AI authoring copilot designed for educational content creation, which uniquely combines a Markov Decision Process (MDP) as a reinforcement learning orchestrator with a locally deployed Qwen3-1.7B-ONNX large language model to iteratively refine text for clarity, unity, and engagement—all running on a modest CPU-only system (Intel i7, 16 GB RAM). Unlike cloud-dependent models, our agent treats writing as a sequential decision problem, selecting refinement actions (e.g., simplification, elaboration) based on real-time LLM and sentiment feedback, ensuring pedagogically sound outputs without internet dependency. Evaluated across five diverse topics, our MDP-orchestrated agent achieved an overall average quality score of 4.23 (on a 0–5 scale), statistically equivalent to leading cloud-based LLMs like ChatGPT and DeepSeek. This performance was validated through blind evaluations by four independent LLMs and human raters, supported by statistical consistency analysis. Our work demonstrates that lightweight local LLMs, when guided by principled MDP policies, can deliver high-quality, context-aware educational content, bridging the gap between powerful AI generation and ethical, on-device deployment. This advancement empowers educators, researchers, and curriculum designers with a trustworthy, accessible tool for intelligent content augmentation aligning with the Quality Education Sustainable Development Goal through innovations in educational technology, inclusive education, equity in education, and lifelong learning.

Paper

Full text

PDF

CPU-Only Self Enhancing Authoring Copilot Design-Based Markov Decision Processes Orchestration and Qwen 3 Local Large Language Model

Semantic Scholar · 2025

Abstract

We introduce a novel, privacy-preserving AI authoring copilot designed for educational content creation, which uniquely combines a Markov Decision Process (MDP) as a reinforcement learning orchestrator with a locally deployed Qwen3-1.7B-ONNX large language model to iteratively refine text for clarity, unity, and engagement—all running on a modest CPU-only system (Intel i7, 16 GB RAM). Unlike cloud-dependent models, our agent treats writing as a sequential decision problem, selecting refinement actions (e.g., simplification, elaboration) based on real-time LLM and sentiment feedback, ensuring pedagogically sound outputs without internet dependency. Evaluated across five diverse topics, our MDP-orchestrated agent achieved an overall average quality score of 4.23 (on a 0–5 scale), statistically equivalent to leading cloud-based LLMs like ChatGPT and DeepSeek. This performance was validated through blind evaluations by four independent LLMs and human raters, supported by statistical consistency analysis. Our work demonstrates that lightweight local LLMs, when guided by principled MDP policies, can deliver high-quality, context-aware educational content, bridging the gap between powerful AI generation and ethical, on-device deployment. This advancement empowers educators, researchers, and curriculum designers with a trustworthy, accessible tool for intelligent content augmentation aligning with the Quality Education Sustainable Development Goal through innovations in educational technology, inclusive education, equity in education, and lifelong learning.

Similar papers

© 2026 NYSGPT2525 LLC