<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>AI | Dylan Chiang</title><link>https://dylanchiang-dev.github.io/en/tags/ai/</link><atom:link href="https://dylanchiang-dev.github.io/en/tags/ai/index.xml" rel="self" type="application/rss+xml"/><description>AI</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-US</language><lastBuildDate>Sun, 24 May 2026 00:00:00 +0000</lastBuildDate><image><url>https://dylanchiang-dev.github.io/media/icon_hu_982c5d63a71b2961.png</url><title>AI</title><link>https://dylanchiang-dev.github.io/en/tags/ai/</link></image><item><title>WeiShi (未識): AI Relationship Exploration for the Sixth Yunnan-Taiwan University Student Innovation and Entrepreneurship Competition</title><link>https://dylanchiang-dev.github.io/en/project/weishi/</link><pubDate>Sun, 24 May 2026 00:00:00 +0000</pubDate><guid>https://dylanchiang-dev.github.io/en/project/weishi/</guid><description>&lt;p>WeiShi is an AI relationship-exploration project developed for the Sixth Yunnan-Taiwan University Student Innovation and Entrepreneurship Competition (第六屆雲台大學生雙創賽). Its proposition is simple: AI avatars connect first, while people retain the final decision. The prototype turns an initial encounter from quick browsing and instant judgment into an AI-assisted process of mutual understanding.&lt;/p>
&lt;h2 id="what-i-am-responsible-for-in-the-project">What I am responsible for in the project&lt;/h2>
&lt;p>I am mainly responsible for product and technology, transforming initial ideas into complete product prototypes that can be actually operated and displayed offline.&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Product Definition&lt;/strong>: Establish the core proposition of &amp;ldquo;knowing before seeing&amp;rdquo; and position AI as a medium for understanding before meeting, rather than a chatbot that makes relationship decisions for users.&lt;/li>
&lt;li>&lt;strong>System Design&lt;/strong>: Planning the complete process of avatar profile creation, Agent Plaza, weekly in-depth recommendations, pre-understanding reports and real-person takeover.&lt;/li>
&lt;li>&lt;strong>Prototype Implementation&lt;/strong>: Connect the main interactions and local status in series to complete a front-end product that can run offline on the competition computer.&lt;/li>
&lt;li>&lt;strong>Governance Implementation&lt;/strong>: Transform the principles of authorization, observability, non-scoring and local priority into actual product nodes and operational restrictions.&lt;/li>
&lt;li>&lt;strong>Result Integration&lt;/strong>: String products, scenarios, governance research and roadshow narratives into a verifiable minimum closed loop, so that innovation is not just a concept.&lt;/li>
&lt;/ul>
&lt;h2 id="project-concept">Project Concept&lt;/h2>
&lt;p>Wei Shi takes &amp;ldquo;knowing before seeing&amp;rdquo; as its core concept. Users first create and continuously calibrate their own AI avatars, and then the two avatars conduct preliminary communications around life goals, relationship expectations, communication methods, and important differences. Only after both parties authorize and confirm their understanding of the report, the system will open the anonymous text conversation with real people.&lt;/p>
&lt;p>The following diagrams are original Chinese-language competition materials retained as historical evidence; the English captions and context on this page are provided for international readers.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="flex justify-center ">
&lt;div class="w-full" >
&lt;img alt="WeiShi five-step pre-understanding mechanism"
srcset="https://dylanchiang-dev.github.io/project/weishi/mechanism_hu_67e2175cb61bcd50.webp 320w, https://dylanchiang-dev.github.io/project/weishi/mechanism_hu_4432c97e62d90d98.webp 480w, https://dylanchiang-dev.github.io/project/weishi/mechanism_hu_efed5a20722e522f.webp 760w"
sizes="(max-width: 480px) 100vw, (max-width: 768px) 90vw, (max-width: 1024px) 80vw, 760px"
src="https://dylanchiang-dev.github.io/project/weishi/mechanism_hu_67e2175cb61bcd50.webp"
width="760"
height="428"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h2 id="core-process">Core process&lt;/h2>
&lt;ol>
&lt;li>&lt;strong>Create AI as me avatar&lt;/strong>: The user corrects memory, values, relationship goals and interaction boundaries.&lt;/li>
&lt;li>&lt;strong>Enter the Agent Square&lt;/strong>: Browse the content shared by other avatars and form a more complete understanding before formal contact.&lt;/li>
&lt;li>&lt;strong>Weekly in-depth recommendation&lt;/strong>: The system provides an in-depth object, presenting the reasons for recommendation, differences to be confirmed, and verification status.&lt;/li>
&lt;li>&lt;strong>Generate pre-understanding report&lt;/strong>: Organize the agreement and differences between the two parties on multiple issues, and retain questions that require real people to answer.&lt;/li>
&lt;li>&lt;strong>Live person takes over the conversation&lt;/strong>: After both parties agree, switch from agent interaction to anonymous text communication.&lt;/li>
&lt;/ol>
&lt;h2 id="core-competitiveness-and-innovation">Core competitiveness and innovation&lt;/h2>
&lt;h3 id="1-from-chat-tools-to-relationship-understanding-protocols">1. From chat tools to relationship understanding protocols&lt;/h3>
&lt;p>Weishi does not pursue more and faster matching, but allows both agents to complete authorizable and reviewable pre-understanding before meeting in person. The value of AI is not to decide relationships for people, but to reduce ineffective chats and information gaps.&lt;/p>
&lt;h3 id="2-in-depth-rhythm-of-one-person-per-week">2. In-depth rhythm of one person per week&lt;/h3>
&lt;p>The product replaces unlimited browsing with &amp;ldquo;one deep subject per week&amp;rdquo;. After multiple rounds of Agent communication, each object forms an understanding report, allowing users to see specific consistencies, differences, and issues to be confirmed.&lt;/p>
&lt;h3 id="3-ai-clone-that-is-observable-and-cannot-exceed-authority">3. AI clone that is observable and cannot exceed authority&lt;/h3>
&lt;p>The avatar can only express within the scope of the user&amp;rsquo;s authorization; the communication process can be reviewed, and sensitive actions need to be confirmed again. The real person always reserves the right to make the final decision.&lt;/p>
&lt;h3 id="4-dont-judge-relationships-by-a-single-score">4. Don’t judge relationships by a single score&lt;/h3>
&lt;p>The understanding report presents evidence, discrepancies and real-life takeover issues on multiple topics, without using an overall score to simplify complex relationships into &amp;ldquo;suitable&amp;rdquo; or &amp;ldquo;unsuitable&amp;rdquo;.&lt;/p>
&lt;h3 id="5-turn-governance-into-a-product-mechanism">5. Turn governance into a product mechanism&lt;/h3>
&lt;p>Local processing, authorized nodes, process traces, and non-repudiation commitments are not additional terms, but product capabilities directly written into the interactive process.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="flex justify-center ">
&lt;div class="w-full" >
&lt;img alt="WeiShi governance framework"
srcset="https://dylanchiang-dev.github.io/project/weishi/governance_hu_ddf37f47d31dfeea.webp 320w, https://dylanchiang-dev.github.io/project/weishi/governance_hu_fdb54d2355e70065.webp 480w, https://dylanchiang-dev.github.io/project/weishi/governance_hu_61deb177996e6346.webp 760w"
sizes="(max-width: 480px) 100vw, (max-width: 768px) 90vw, (max-width: 1024px) 80vw, 760px"
src="https://dylanchiang-dev.github.io/project/weishi/governance_hu_ddf37f47d31dfeea.webp"
width="760"
height="428"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h2 id="design-principles">Design principles&lt;/h2>
&lt;ul>
&lt;li>&lt;strong>Human-in-the-loop&lt;/strong>: AI assists in organizing and understanding, but does not make promises, define relationships, or decide to meet for the user.&lt;/li>
&lt;li>&lt;strong>Consent takes priority&lt;/strong>: clear authorization nodes are reserved for pre-contact, report confirmation and real-person takeover.&lt;/li>
&lt;li>&lt;strong>Local First&lt;/strong>: Competition prototypes can be run offline, and personal correction status is saved locally in the browser.&lt;/li>
&lt;li>&lt;strong>Understand not score&lt;/strong>: The report presents specific issues and evidence, and does not simplify the relationship with a single matching score.&lt;/li>
&lt;/ul>
&lt;p>The current version is a local front-end prototype used for competition display and product verification. It does not include a formal account, online model service, real matching algorithm or production data platform.&lt;/p></description></item><item><title>Report Reading: Generative Artificial Intelligence Application Development Report (2025)</title><link>https://dylanchiang-dev.github.io/en/post/gen-ai-report-2025-reading/</link><pubDate>Fri, 26 Dec 2025 00:00:00 +0000</pubDate><guid>https://dylanchiang-dev.github.io/en/post/gen-ai-report-2025-reading/</guid><description>&lt;h1 id="report-information">Report information&lt;/h1>
&lt;ul>
&lt;li>&lt;strong>Title&lt;/strong>: Generative Artificial Intelligence Application Development Report (2025)&lt;/li>
&lt;li>&lt;strong>Issuing agency&lt;/strong>: China Internet Network Information Center (CNNIC)&lt;/li>
&lt;li>&lt;strong>Release Date&lt;/strong>: October 2025&lt;/li>
&lt;li>&lt;strong>Original link&lt;/strong>: [Generative Artificial Intelligence Application Development Report (2025)] (
) (Note: This is CNNIC’s typical reporting path, please refer to the official website for details)&lt;/li>
&lt;li>&lt;strong>Core Content&lt;/strong>: Based on the framework of &amp;ldquo;User Popularization-Industrial Development-Typical Applications-Development Environment&amp;rdquo;, analyze the development status and future prospects of generative AI.&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h2 id="comprehensive-overview-and-summary-of-the-report">Comprehensive overview and summary of the report&lt;/h2>
&lt;p>This report records in detail the explosive growth of generative artificial intelligence (GenAI) in 2025, especially the rapid popularity and technological iteration of the Chinese market.&lt;/p>
&lt;h3 id="1-development-characteristics-domestic-achievements-and-efficiency-breakthroughs">1. Development characteristics: domestic achievements and efficiency breakthroughs&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Scale Explosion&lt;/strong>: As of June 2025, the number of GenAI users in China reached &lt;strong>515 million&lt;/strong>, with a penetration rate of &lt;strong>36.5%&lt;/strong>.&lt;/li>
&lt;li>&lt;strong>The Rise of DeepSeek&lt;/strong>: DeepSeek-R1 achieves excellent performance at less than 1/10 the cost of similar models, breaking the &amp;ldquo;computing power determinism&amp;rdquo; and topping the application list in 140 countries around the world.&lt;/li>
&lt;li>&lt;strong>Technical Trends&lt;/strong>: Logical reasoning capabilities have been significantly improved, multi-modal (Venture Video, Tusheng Audio) has developed by leaps and bounds, reasoning costs have been significantly reduced, and lightweight models (device-side AI) empower more terminal devices.&lt;/li>
&lt;/ul>
&lt;h3 id="2-user-popularity">2. User popularity&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Market structure&lt;/strong>: &lt;strong>Doubao&lt;/strong> and &lt;strong>DeepSeek&lt;/strong> occupy the leading position in the market.&lt;/li>
&lt;li>&lt;strong>Application motivation&lt;/strong>: Mainly used for &lt;strong>answering questions (80.9%)&lt;/strong>, &lt;strong>text processing (36.0%)&lt;/strong> and &lt;strong>audio and video generation (33.0%)&lt;/strong>.&lt;/li>
&lt;li>&lt;strong>Group Characteristics&lt;/strong>: Young and middle-aged users are the main force, with users aged 19 and under accounting for 33.8%, showing the young generation’s high acceptance of AI technology.&lt;/li>
&lt;/ul>
&lt;h3 id="3-typical-application-scenarios">3. Typical application scenarios&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Industry and Agriculture&lt;/strong>: From smart irrigation, precision agriculture to industrial robots (such as UBTECH Walker S1) in production line training, AI is reshaping the labor structure.&lt;/li>
&lt;li>&lt;strong>Life Services&lt;/strong>:
&lt;ul>
&lt;li>&lt;strong>Smart Search&lt;/strong>: Shift from &amp;ldquo;finding links&amp;rdquo; to &amp;ldquo;getting answers&amp;rdquo;, Search as a Service.&lt;/li>
&lt;li>&lt;strong>Content Creation&lt;/strong>: Sora and Keling AI bring short drama and advertising production into the era of &amp;ldquo;second-level generation&amp;rdquo;.&lt;/li>
&lt;li>&lt;strong>Office Assistant&lt;/strong>: AI code generation (more than 30% of new code is generated by AI) and intelligent document processing become the norm.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="4-development-environment-and-prospects">4. Development environment and prospects&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Policy Support&lt;/strong>: Establish an artificial intelligence security governance framework and promote &amp;ldquo;artificial intelligence +&amp;rdquo; actions.&lt;/li>
&lt;li>&lt;strong>Future Trends&lt;/strong>: &lt;strong>Agent&lt;/strong> has autonomous decision-making and execution capabilities, &lt;strong>Embodied Intelligence&lt;/strong> allows robots to enter the physical world and interact, &lt;strong>Scientific Research (AI for Science)&lt;/strong> accelerates lunar research and weather prediction.&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h2 id="my-understanding">My understanding&lt;/h2>
&lt;p>This report is not only a accumulation of technical data, but also reveals the paradigm shift of AI technology from &amp;ldquo;concept laboratory&amp;rdquo; to &amp;ldquo;full employee productivity&amp;rdquo;.&lt;/p>
&lt;h3 id="1-implications-for-personal-research-fields-political-work-and-assistant-behavior">1. Implications for personal research fields (political work and assistant behavior)&lt;/h3>
&lt;p>The report pointed out that &lt;strong>80.9% of users’ primary need is to “answer questions”&lt;/strong>, and in the office assistant scenario, AI’s understanding and polishing capabilities have become the core. For my study of Legislative Assistants’ Use of Generative AI, this provides strong external validity:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Efficiency Liberated&lt;/strong>: Political assistants are faced with the collection of data on a large number of public issues and the preparation of press releases. The high usage of AI search and document processing indicates that there will be an &amp;ldquo;automated revolution&amp;rdquo; in political work processes.&lt;/li>
&lt;li>&lt;strong>Decision Assistance&lt;/strong>: The &amp;ldquo;improvement in reasoning capabilities&amp;rdquo; mentioned in the report means that the assistant will not only use AI to write drafts in the future, but may also use AI to conduct policy impact assessment and voter sentiment analysis.&lt;/li>
&lt;/ul>
&lt;h3 id="2-inspiration-for-the-development-of-ai-in-taiwan-and-cross-strait">2. Inspiration for the development of AI in Taiwan and cross-strait&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Lessons from cost advantages&lt;/strong>: The success of DeepSeek shows that Taiwan, with limited computing resources, should focus on &lt;strong>algorithm optimization and deep cultivation of specific vertical fields (Vertical AI)&lt;/strong> instead of simply pursuing parameter scale.&lt;/li>
&lt;li>&lt;strong>Complementarity of application scenarios&lt;/strong>: Mainland China has a large scale of industrial and agricultural applications, while Taiwan has advantages in &lt;strong>semiconductor end-side AI&lt;/strong> and &lt;strong>high-quality service trade&lt;/strong>. Both sides of the Taiwan Strait face common challenges in AI governance (such as copyright, data privacy), and there is room for cross-regional dialogue.&lt;/li>
&lt;/ul>
&lt;h3 id="3-thoughts-on-the-direction-of-academic-research">3. Thoughts on the direction of academic research&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>From &amp;ldquo;use&amp;rdquo; to &amp;ldquo;collaboration&amp;rdquo;&lt;/strong>: Future academic research should not only focus on whether assistants use AI, but should study how the &amp;ldquo;human-machine collaboration model&amp;rdquo; changes the power structure in the political field.&lt;/li>
&lt;li>&lt;strong>Update of Technology Acceptance Model&lt;/strong>: The traditional TAM model may not be enough to explain the changes brought about by Agentic workflow. Research needs to focus on users&amp;rsquo; psychological boundaries and ethical trust in AI &amp;ldquo;autonomy&amp;rdquo;.&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>The &amp;ldquo;2025 Report&amp;rdquo; marks that AI has entered the &amp;ldquo;full-scenario penetration&amp;rdquo; stage from a single breakthrough. As researchers, we need to pay close attention to the dynamic balance between the &amp;ldquo;efficiency dividend&amp;rdquo; and &amp;ldquo;ethical risks&amp;rdquo; brought by AI, especially in the highly sensitive field of political work.&lt;/p></description></item><item><title>Report reading: Microsoft AI Diffusion Report (2025)</title><link>https://dylanchiang-dev.github.io/en/post/microsoft-ai-diffusion-report-2025-reading/</link><pubDate>Fri, 26 Dec 2025 00:00:00 +0000</pubDate><guid>https://dylanchiang-dev.github.io/en/post/microsoft-ai-diffusion-report-2025-reading/</guid><description>&lt;h1 id="report-information">Report information&lt;/h1>
&lt;ul>
&lt;li>&lt;strong>Title&lt;/strong>: Microsoft AI Diffusion Report: Mapping Global AI Adoption and Innovation&lt;/li>
&lt;li>&lt;strong>Published by&lt;/strong>: Microsoft Research&lt;/li>
&lt;li>&lt;strong>Release Date&lt;/strong>: October 2025&lt;/li>
&lt;li>&lt;strong>Original link&lt;/strong>:
&lt;/li>
&lt;li>&lt;strong>Core content&lt;/strong>: Track the global adoption status, infrastructure distribution, narrowing of the technology frontier and challenges faced by AI as the fastest spreading technology in history.&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h2 id="comprehensive-overview-and-summary-of-the-report">Comprehensive overview and summary of the report&lt;/h2>
&lt;p>This Microsoft report focuses on the &amp;ldquo;diffusion&amp;rdquo; dynamics of AI technology on a global scale, revealing the contradictory current situation of technology popularization and resource concentration coexisting.&lt;/p>
&lt;h3 id="1-the-fastest-technological-diffusion-in-history">1. The fastest technological diffusion in history&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Breakthrough Velocity&lt;/strong>: AI has attracted more than &lt;strong>1.2 billion users&lt;/strong> in less than three years. Its diffusion rate is far faster than any previous general purpose technology (GPT), such as the Internet, personal computers or smartphones.&lt;/li>
&lt;li>&lt;strong>Adoption rate differentiation&lt;/strong>: Although the adoption rate is fast, there are significant differences between the Global North vs. Global South. The adoption rate in northern countries is approximately &lt;strong>2 times&lt;/strong> that of the South.&lt;/li>
&lt;/ul>
&lt;h3 id="2-frontier-trends-in-infrastructure-and-models">2. Frontier trends in infrastructure and models&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Concentration of computing power and data&lt;/strong>: &lt;strong>86% of the world’s data center capacity&lt;/strong> is concentrated in the United States and China. This extreme concentration of infrastructure creates challenges for the democratization of AI.&lt;/li>
&lt;li>&lt;strong>Frontier Narrowing&lt;/strong>: Although the United States (such as OpenAI&amp;rsquo;s GPT-5 and other benchmarks) still maintains the lead, the performance gap of latecomers is rapidly narrowing. The report states that China is estimated to be less than &lt;strong>6 months&lt;/strong> behind on the technological frontier.&lt;/li>
&lt;li>&lt;strong>Model Distribution&lt;/strong>: The world&amp;rsquo;s top 200 models are only concentrated in 7 countries (the United States, China, France, South Korea, the United Kingdom, Canada, and Israel), showing the high threshold for technology research and development.&lt;/li>
&lt;/ul>
&lt;h3 id="3-diffusion-leaders">3. Diffusion Leaders&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Non-R&amp;amp;D adoption leaders&lt;/strong>: Countries such as the United Arab Emirates (59.4%), Singapore (58.6%), and Norway (45.3%) stand out.&lt;/li>
&lt;li>&lt;strong>Success Factors&lt;/strong>: These countries have proven that even without developing underlying models, they can maintain global leadership in AI adoption through &lt;strong>strong education systems, friendly policy environments, and digital infrastructure&lt;/strong>.&lt;/li>
&lt;/ul>
&lt;h3 id="4-key-barriers-language-and-infrastructure">4. Key barriers: language and infrastructure&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Language inequality&lt;/strong>: Adoption rates are significantly lower in low-resource language areas (e.g. Malawi, Laos). The language tolerance of AI models has become a key bottleneck for popularization.&lt;/li>
&lt;li>&lt;strong>Infrastructure constraints&lt;/strong>: About half of the world’s population (4 billion people) still lack the basic conditions (electricity, network and broadband) required to use AI.&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h2 id="my-understanding">My understanding&lt;/h2>
&lt;p>Microsoft’s global perspective provides a broader context for understanding AI’s place in politics, bipartisanship, and academia.&lt;/p>
&lt;h3 id="1-implications-for-personal-research-fields-political-work-and-assistant-behavior">1. Implications for personal research fields (political work and assistant behavior)&lt;/h3>
&lt;p>The report mentioned that the rapid spread of AI indicates that the &amp;ldquo;digital generation gap&amp;rdquo; in the political field will be quickly compressed:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Global knowledge equality and challenges&lt;/strong>: If 1.2 billion users are using AI, political assistants are no longer a question of &amp;ldquo;whether to use them&amp;rdquo;, but a question of how to maintain &amp;ldquo;political uniqueness&amp;rdquo; during use.&lt;/li>
&lt;li>&lt;strong>Skills first&lt;/strong>: The cases of the United Arab Emirates and Singapore illustrate that &amp;ldquo;skills use&amp;rdquo; and &amp;ldquo;policy guidance&amp;rdquo; have a more direct impact on the application side than &amp;ldquo;model development&amp;rdquo;. This supports the hypothesis in my research that focuses on assistants’ personal technology preferences and environmental support.&lt;/li>
&lt;/ul>
&lt;h3 id="2-inspiration-for-the-development-of-ai-in-taiwan-and-cross-strait">2. Inspiration for the development of AI in Taiwan and cross-strait&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Strategic opportunities for frontier narrowing&lt;/strong>: When the frontier technology gap narrows to within 6 months, Taiwan, as the core of the global supply chain, will have a greater say in hardware support in the &amp;ldquo;inference&amp;rdquo; stage.&lt;/li>
&lt;li>&lt;strong>Learn from the &amp;ldquo;Adoption Leader&amp;rdquo; model&lt;/strong>: Taiwan should become a world-leading &amp;ldquo;AI efficient adoption zone&amp;rdquo; through regulatory innovation and education transformation, like Singapore or Norway, despite the lack of local ultra-large-scale underlying models.&lt;/li>
&lt;/ul>
&lt;h3 id="3-thoughts-on-the-direction-of-academic-research">3. Thoughts on the direction of academic research&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Focus on the imbalance of &amp;ldquo;AI diffusion&amp;rdquo;&lt;/strong>: Research should not be limited to advanced regions, but should focus on whether AI has exacerbated the gap between disadvantaged groups (or small parties, resource-poor politicians) and resource concentrators.&lt;/li>
&lt;li>&lt;strong>Language specificity research&lt;/strong>: For Taiwan, which uses traditional Chinese, we should study how &amp;ldquo;language specificity&amp;rdquo; affects the model&amp;rsquo;s understanding and expression accuracy in the local political context.&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>If CNNIC&amp;rsquo;s report shows the &amp;ldquo;depth&amp;rdquo; of China&amp;rsquo;s AI applications, then Microsoft&amp;rsquo;s report shows the &amp;ldquo;breadth&amp;rdquo; and &amp;ldquo;imbalance&amp;rdquo; on a global scale. The future of AI depends not only on who has the strongest model, but also on who can harness this power the fastest and fairest.&lt;/p></description></item><item><title>Report reading: OpenRouter State of AI (2025)</title><link>https://dylanchiang-dev.github.io/en/post/openrouter-state-of-ai-2025-reading/</link><pubDate>Wed, 10 Dec 2025 00:00:00 +0000</pubDate><guid>https://dylanchiang-dev.github.io/en/post/openrouter-state-of-ai-2025-reading/</guid><description>&lt;h1 id="report-information">Report information&lt;/h1>
&lt;ul>
&lt;li>&lt;strong>Title&lt;/strong>: State of AI | OpenRouter (Empirical Study)&lt;/li>
&lt;li>&lt;strong>Publishing Authority&lt;/strong>: OpenRouter&lt;/li>
&lt;li>&lt;strong>Published&lt;/strong>: December 5, 2024 (reported observation period spans 2024-2025)&lt;/li>
&lt;li>&lt;strong>Original link&lt;/strong>:
&lt;/li>
&lt;li>&lt;strong>Data Foundation&lt;/strong>: Real interactive metadata based on OpenRouter unified inference layer &lt;strong>100 trillion (Trillion) tokens&lt;/strong>, covering 300+ models and 60+ providers around the world.&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h2 id="report-key-insights-abstract--empirical-findings">Report Key Insights (Abstract &amp;amp; Empirical Findings)&lt;/h2>
&lt;p>OpenRouter&amp;rsquo;s report avoids traditional subjective evaluations and instead starts from the &amp;ldquo;real behavior&amp;rdquo; of large-scale production environments, revealing in-depth patterns in the following dimensions:&lt;/p>
&lt;h3 id="1-paradigm-shift-of-agentic-inference">1. Paradigm shift of Agentic Inference&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Absolute dominance of inference models&lt;/strong>: The report points out that since the release of the o1 class model at the end of 2024, AI usage has shifted from &amp;ldquo;content generation&amp;rdquo; to &amp;ldquo;multi-step reasoning.&amp;rdquo; By 2025, more than 50% of total token traffic will flow to inference optimization models (led by xAI’s Grok Code Fast, followed by the Gemini 2.5 series and DeepSeek R1).&lt;/li>
&lt;li>&lt;strong>Tool-Calling Trend&lt;/strong>: Data display tool calling is no longer an option for developers, but a default for high-value workflows. The Claude 3.5/3.7 series dominated in the early days, and then Grok and GLM 4.5 quickly entered the market, reflecting that &lt;strong>action through planning&lt;/strong> is the moat for future models.&lt;/li>
&lt;/ul>
&lt;h3 id="2-dynamic-balance-between-open-source-and-closed-source-market-equilibrium">2. Dynamic balance between open source and closed source (Market Equilibrium)&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>30% &amp;ldquo;Open Source Ceiling&amp;rdquo;&lt;/strong>: Although closed source models still account for 70% of the market share (focusing on regulated enterprise-level workflows), open source/weighted open models (OSS) have stabilized at around 30%.&lt;/li>
&lt;li>&lt;strong>The Rise of China’s Open Source Model&lt;/strong>: The release of DeepSeek V3 and Qwen 3 Coder directly led to a surge in usage. Especially in the field of code assistance (Programming), the Chinese open source model briefly accounted for more than half of OSS code tasks in mid-2025.&lt;/li>
&lt;/ul>
&lt;h3 id="3-model-family-usage-profiles-provider-profiles">3. Model family usage profiles (Provider Profiles)&lt;/h3>
&lt;p>The report reveals users’ “cognitive division of labor” towards different model brands:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Anthropic (Claude)&lt;/strong>: Extreme focus on &lt;strong>Programming and Technology (80%+)&lt;/strong>. Users consider Claude their go-to choice for complex reasoning and engineering.&lt;/li>
&lt;li>&lt;strong>Google (Gemini)&lt;/strong>: The most diverse performance, covering translation, science, law and general knowledge, showing the characteristics of &lt;strong>&amp;ldquo;digital encyclopedia/information engine&amp;rdquo;&lt;/strong>.&lt;/li>
&lt;li>&lt;strong>OpenAI (GPT)&lt;/strong>: It has undergone a transformation from early science and general knowledge to deeper developer workflow and productivity tools. Its positioning is between Claude&amp;rsquo;s professionalism and Google&amp;rsquo;s diversity.&lt;/li>
&lt;li>&lt;strong>DeepSeek&lt;/strong>: shows an amazing &lt;strong>consumer-heavy&lt;/strong>, with more than 2/3 of the traffic coming from creativity, entertainment and role-playing.&lt;/li>
&lt;/ul>
&lt;h3 id="4-retention-analysis-cinderellas-glass-slipper-effect">4. Retention Analysis: Cinderella’s “Glass Slipper Effect”&lt;/h3>
&lt;p>This is the most interesting finding in the report. The study observed:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Foundational Cohorts&lt;/strong>: Early users (such as first-month users of Gemini 2.5 or Claude 4 Sonnet) have retention rates of up to 40% after 5 months, much higher than those of subsequent users.&lt;/li>
&lt;li>&lt;strong>First solution advantage&lt;/strong>: When a model is the first to solve a specific problem (such as a complex Tool-use or logic difficulty), the user group will have a strong &lt;strong>path dependence (Cognitive inertia)&lt;/strong>. This is the &amp;ldquo;glass shoe&amp;rdquo; effect: once adapted, a powerful locking effect will occur.&lt;/li>
&lt;/ul>
&lt;h3 id="5-cost-and-demand-elasticity-jevons-paradox">5. Cost and Demand Elasticity (Jevons Paradox)&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Price Inelastic&lt;/strong>: Interestingly, a 10% price drop only drives 0.5-0.7% usage growth. This shows that &amp;ldquo;quality and trust&amp;rdquo; are far more important than price, and top companies are willing to pay a premium for stability.&lt;/li>
&lt;li>&lt;strong>Jevons Paradox&lt;/strong>: Although the price elasticity is low at the macro level, in the field of &amp;ldquo;Efficient Giants&amp;rdquo; (such as Gemini Flash or DeepSeek), low cost does induce larger Token consumption (users start to perform more iterations and longer context queries).&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h2 id="my-understanding-academic-and-practical-inspiration-behind-the-data">My understanding: academic and practical inspiration behind the data&lt;/h2>
&lt;p>This &amp;ldquo;empirical&amp;rdquo; report elevates the discussion of AI from &amp;ldquo;whether it is easy to use&amp;rdquo; to &amp;ldquo;how to deploy it systematically.&amp;rdquo;&lt;/p>
&lt;h3 id="1-implications-for-personal-research-fields-political-work-and-assistant-behavior">1. Implications for personal research fields (political work and assistant behavior)&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Inference Dividends and Agentization&lt;/strong>: More than 50% of the models in the report turned to inference, validating my observation that assistants in political work are leveraging AI to handle more &amp;ldquo;judgmental&amp;rdquo; tasks. Future research should shift from &amp;ldquo;whether assistants use AI&amp;rdquo; to &amp;ldquo;how assistants guide public opinion through AI&amp;rsquo;s agentic workflow.&amp;rdquo;&lt;/li>
&lt;li>&lt;strong>Scenario Adaptation (Glass Slipper)&lt;/strong>: For political assistants, the first model that can accurately craft their own tone or analyze the sensitivities of a constituency will create loyalties that are extremely difficult to break. This explains why some offices are stuck on older versions of GPT or specific models.&lt;/li>
&lt;/ul>
&lt;h3 id="2-inspiration-for-the-development-of-ai-in-taiwan-and-cross-strait">2. Inspiration for the development of AI in Taiwan and cross-strait&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Export advantages of open source models&lt;/strong>: The explosive growth of DeepSeek and Qwen on OpenRouter proves that &amp;ldquo;open source orientation&amp;rdquo; is the main driver of Chinese models going global. Taiwan can boldly integrate these &amp;ldquo;Efficient Giants&amp;rdquo; on the application side and use cost savings for front-end scene optimization.&lt;/li>
&lt;li>&lt;strong>The necessity of multi-model architecture&lt;/strong>: Since no single model can dominate all portraits (DeepSeek for Roleplay, Claude for Code), Taiwanese companies and think tanks should adopt &lt;strong>&amp;ldquo;Multi-model Stack&amp;rdquo;&lt;/strong> to achieve the best balance between cost and performance.&lt;/li>
&lt;/ul>
&lt;h3 id="3-thoughts-on-the-direction-of-academic-research">3. Thoughts on the direction of academic research&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Focus on the bias of agent inference&lt;/strong>: When AI begins to autonomously call tools and plan paths (Agentic Inference), its bias will no longer just be &amp;ldquo;saying the wrong thing&amp;rdquo;, but &amp;ldquo;doing the wrong thing&amp;rdquo;. This is an extremely critical new topic in political and legal studies.&lt;/li>
&lt;li>&lt;strong>Application of Retention Rate Research&lt;/strong>: We should study how to shorten the journey for users to find the &amp;ldquo;glass slipper&amp;rdquo; model to increase the success rate of digital transformation.&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>OpenRouter&amp;rsquo;s data reveals a cruel but hopeful truth: In the great era of AI, the emergence of inference models such as o1 has not capped competition, but has opened up a new battlefield of &amp;ldquo;multi-step operation&amp;rdquo; and &amp;ldquo;scene adaptation&amp;rdquo;**. The future does not belong to the person with the biggest model, but to the person who can accurately find the &amp;ldquo;glass slipper&amp;rdquo;.&lt;/p></description></item><item><title>Paper reading: How People Use ChatGPT - In-depth analysis of the ChatGPT usage behavior of 700 million users around the world</title><link>https://dylanchiang-dev.github.io/en/post/chatgpt-usage-economics/</link><pubDate>Tue, 04 Nov 2025 00:00:00 +0000</pubDate><guid>https://dylanchiang-dev.github.io/en/post/chatgpt-usage-economics/</guid><description>&lt;p>I recently read the important research paper &amp;ldquo;How People Use ChatGPT&amp;rdquo; from the research team of OpenAI, Duke University and Harvard University. This is the first large-scale usage behavior study based on internal data of ChatGPT. Through innovative privacy protection methods, the study analyzed 26 billion messages from 700 million users from the launch of ChatGPT in November 2022 to July 2025, revealing the actual usage patterns and economic value of generative AI.&lt;/p>
&lt;h2 id="research-methods-and-data">Research methods and data&lt;/h2>
&lt;h3 id="privacy-protecting-automated-classification-system">Privacy-protecting automated classification system&lt;/h3>
&lt;p>The biggest technical highlight of this research is its privacy protection method:&lt;/p>
&lt;p>&lt;strong>Automated classification process&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Automatically analyze message content using LLM classifier, humans never view the original message&lt;/li>
&lt;li>First remove sensitive information through PII cleaning tools&lt;/li>
&lt;li>Only aggregated results are analyzed, any query must return a combination of at least 100 users&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Classification Category&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>WORK/NON-WORK USE&lt;/strong>: Based on whether the message is related to paid work&lt;/li>
&lt;li>&lt;strong>Conversation Topics&lt;/strong>: 24 subcategories, summarized into 7 major themes&lt;/li>
&lt;li>&lt;strong>Interaction intent&lt;/strong>: Asking, Doing, Expressing&lt;/li>
&lt;li>&lt;strong>WORK ACTIVITIES&lt;/strong>: 332 intermediate-level work activities based on O*NET system&lt;/li>
&lt;/ul>
&lt;h3 id="data-sample">Data sample&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Main Sample&lt;/strong>: Random sample of 1.1 million conversations from May 2024 to June 2025&lt;/li>
&lt;li>&lt;strong>User Sample&lt;/strong>: A subset of approximately 130,000 users used for demographic analysis&lt;/li>
&lt;li>&lt;strong>Exclusion Conditions&lt;/strong>: Users who have not logged in, users under 18 years old, users who have deleted their accounts, and users who have opted out of training&lt;/li>
&lt;/ul>
&lt;h2 id="-1-growth-and-structure-explosive-growth-of-non-work-purposes">📈 1. Growth and structure: explosive growth of non-work purposes&lt;/h2>
&lt;h3 id="overall-growth-data">Overall growth data&lt;/h3>
&lt;p>&lt;strong>User size&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>July 2025: &lt;strong>700 million weekly active users&lt;/strong> (approximately 10% of the global adult population)&lt;/li>
&lt;li>Average daily message volume: &lt;strong>2.5 billion&lt;/strong> (29,000 messages per second)&lt;/li>
&lt;li>Growth rate: The fastest spreading technology in history, surpassing all precedents&lt;/li>
&lt;/ul>
&lt;h3 id="non-work-usage-increases-faster">Non-work usage increases faster&lt;/h3>
&lt;p>&lt;strong>Core Finding&lt;/strong>: Non-work-related uses are growing much faster than work uses.&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Time point&lt;/th>
&lt;th>Non-work messages&lt;/th>
&lt;th>Proportion&lt;/th>
&lt;th>Work messages&lt;/th>
&lt;th>Proportion&lt;/th>
&lt;th>Total&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>June 2024&lt;/td>
&lt;td>238 million&lt;/td>
&lt;td>53%&lt;/td>
&lt;td>213 million&lt;/td>
&lt;td>47%&lt;/td>
&lt;td>451 million&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>June 2025&lt;/td>
&lt;td>1.911 billion&lt;/td>
&lt;td>73%&lt;/td>
&lt;td>716 million&lt;/td>
&lt;td>27%&lt;/td>
&lt;td>2.627 billion&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>&lt;strong>Key Insights&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Non-work messages increased &lt;strong>8 times&lt;/strong> (238%), work messages increased &lt;strong>3.4 times&lt;/strong> (236%)&lt;/li>
&lt;li>In June 2025, non-work use accounted for &lt;strong>73%&lt;/strong>, which is absolutely dominant&lt;/li>
&lt;li>This change mainly comes from changes in the usage patterns of existing users rather than changes in the composition of new users&lt;/li>
&lt;/ul>
&lt;h3 id="use-dynamic-evolution-of-topics">Use dynamic evolution of topics&lt;/h3>
&lt;p>&lt;strong>Three mainstream uses&lt;/strong> (accounting for nearly 80% of total use):&lt;/p>
&lt;ol>
&lt;li>
&lt;p>&lt;strong>Practical Guidance&lt;/strong>: long-term stability at about 29%&lt;/p>
&lt;ul>
&lt;li>Tutorial teaching (accounting for 36% of practical guidelines)&lt;/li>
&lt;li>How-to suggestions (accounting for 30% of practical guidance)&lt;/li>
&lt;li>Creative ideas&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Writing&lt;/strong>: 36% → 24% (decline within one year)&lt;/p>
&lt;ul>
&lt;li>But it is still the &lt;strong>first largest category&lt;/strong> in work use (about 40%)&lt;/li>
&lt;li>The management/business group has a higher usage ratio (&amp;gt;50%)&lt;/li>
&lt;li>&lt;strong>Key findings&lt;/strong>: About 2/3 of the writing uses are to modify the text provided by users (editing, criticizing, translating, summarizing) rather than creating from scratch&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Seeking Information&lt;/strong>: 14% → 24% (rapid increase)&lt;/p>
&lt;ul>
&lt;li>Search for specific people, events, products, recipes and more&lt;/li>
&lt;li>Become a &lt;strong>closer alternative&lt;/strong> to web search&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ol>
&lt;p>&lt;strong>Other theme variations&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>
&lt;p>&lt;strong>Technical Help&lt;/strong>: 12% → ~5%&lt;/p>
&lt;ul>
&lt;li>Programming related accounted for only 4.2%, significantly lower than expected&lt;/li>
&lt;li>May switch to IDE plug-ins, professional programming tools or API scenarios&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Multimedia&lt;/strong>: 2% → &amp;gt;7%&lt;/p>
&lt;ul>
&lt;li>Short-term jump after the image generation function is launched in April 2025&lt;/li>
&lt;li>Subsequent pullback but maintains higher baseline&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="-2-work-scenarios-and-tasks-ai-as-a-decision-support-system">💼 2. Work scenarios and tasks: AI as a decision support system&lt;/h2>
&lt;h3 id="writing-the-common-mother-task-of-white-collar-workers">Writing: The common mother task of white-collar workers&lt;/h3>
&lt;p>Among work-related messages, &lt;strong>writing accounts for about 40%&lt;/strong> and is the most important work purpose:&lt;/p>
&lt;p>&lt;strong>Career Differences&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Management/Business: &lt;strong>52%&lt;/strong> of work-related news is writing&lt;/li>
&lt;li>Education/Medical: &lt;strong>49-50%&lt;/strong>&lt;/li>
&lt;li>Computer related: &lt;strong>Relatively low&lt;/strong>, more emphasis on technical assistance&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Content Analysis&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>About &lt;strong>2/3&lt;/strong> of writing requests are to revise existing text (editing, criticizing, translating, summarizing)&lt;/li>
&lt;li>About &lt;strong>1/3&lt;/strong> is created from scratch (new emails, briefings, proposals, etc.)&lt;/li>
&lt;li>This explains the high satisfaction and steady growth of writing applications: the risks are manageable and can be directly integrated into existing workflows&lt;/li>
&lt;/ul>
&lt;h3 id="work-activity-analysis-based-on-onet">Work activity analysis based on O*NET&lt;/h3>
&lt;p>A study mapping work messages to the U.S. Department of Labor’s O*NET work activity system found:&lt;/p>
&lt;p>&lt;strong>Seven major work activities cover approximately 77% of all messages&lt;/strong>:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Activity Categories&lt;/th>
&lt;th>All News&lt;/th>
&lt;th>Work News&lt;/th>
&lt;th>Features&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Get information&lt;/td>
&lt;td>19.3%&lt;/td>
&lt;td>6.7%&lt;/td>
&lt;td>Focus more on professional information in work settings&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Interpret information for others&lt;/td>
&lt;td>13.1%&lt;/td>
&lt;td>7.3%&lt;/td>
&lt;td>Collaboration and knowledge transfer&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Record/documented information&lt;/td>
&lt;td>12.8%&lt;/td>
&lt;td>13.2%&lt;/td>
&lt;td>&lt;strong>The first category of work scenarios&lt;/strong>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Providing consultation and advice&lt;/td>
&lt;td>9.2%&lt;/td>
&lt;td>3.1%&lt;/td>
&lt;td>Professional service core&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Creative thinking&lt;/td>
&lt;td>9.1%&lt;/td>
&lt;td>9.3%&lt;/td>
&lt;td>Problem solving and innovation&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Decision-making and problem-solving&lt;/td>
&lt;td>8.5%&lt;/td>
&lt;td>10.6%&lt;/td>
&lt;td>&lt;strong>The second largest category of work scenarios&lt;/strong>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Working with computers&lt;/td>
&lt;td>4.9%&lt;/td>
&lt;td>7.7%&lt;/td>
&lt;td>Technology-intensive jobs&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>&lt;strong>High homogeneity across occupations&lt;/strong>:&lt;/p>
&lt;p>Regardless of management, engineering, education, medical or administrative occupations, the &lt;strong>top 5 work activities are almost the same&lt;/strong>:&lt;/p>
&lt;ol>
&lt;li>Obtain information&lt;/li>
&lt;li>Decision-making and problem-solving&lt;/li>
&lt;li>Record/Documentation&lt;/li>
&lt;li>Think creatively&lt;/li>
&lt;li>Explain information to others&lt;/li>
&lt;/ol>
&lt;p>This shows that ChatGPT’s value creation model in different occupations is highly consistent.&lt;/p>
&lt;h2 id="-3-interaction-type-and-experience-transformation-from-execution-to-thinking">🎯 3. Interaction type and experience: transformation from execution to thinking&lt;/h2>
&lt;h3 id="askingdoingexpressing-framework">Asking/Doing/Expressing Framework&lt;/h3>
&lt;p>The study divided user intent into three categories and found significant trend changes:&lt;/p>
&lt;p>&lt;strong>Overall Distribution&lt;/strong> (May 2024):&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Asking&lt;/strong>: 49% - Seeking information or advice to help make decisions&lt;/li>
&lt;li>&lt;strong>Doing&lt;/strong>: 40% - Request to complete a specific task&lt;/li>
&lt;li>&lt;strong>Expressing&lt;/strong>: 11% - Expressing opinions or feelings&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Trend Change&lt;/strong> (to June 2025):&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Asking&lt;/strong>: 51.6% (↑)&lt;/li>
&lt;li>&lt;strong>Doing&lt;/strong>: 34.6% (↓)&lt;/li>
&lt;li>&lt;strong>Expressing&lt;/strong>: 13.8% (↑)&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Differences in work scenarios&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Asking: 35%&lt;/li>
&lt;li>Doing: 56% (about 75% is writing tasks)&lt;/li>
&lt;li>Expressing: 9%&lt;/li>
&lt;/ul>
&lt;h3 id="experience-quality-analysis">Experience quality analysis&lt;/h3>
&lt;p>&lt;strong>Overall Satisfaction Growth&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Positive/negative review ratio: from about &lt;strong>3:1&lt;/strong> → &lt;strong>4:1&lt;/strong>&lt;/li>
&lt;li>Experience quality is highly related to usage intention&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Positive rating by topic&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Self-expression&lt;/strong>: Highest (good/bad ratio &amp;gt;7)&lt;/li>
&lt;li>&lt;strong>Multimedia&lt;/strong>: Lower (about 1.7)&lt;/li>
&lt;li>&lt;strong>Technical Help&lt;/strong>: Low (~2.7)&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Classification by Intent&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>The satisfaction level of &lt;strong>Asking&lt;/strong> is significantly higher than that of Doing and Expressing&lt;/li>
&lt;li>This is consistent with the core value of &amp;ldquo;helping thinking and decision-making&amp;rdquo;&lt;/li>
&lt;/ul>
&lt;h2 id="-4-ethnicity-and-diffusion-from-elite-tools-to-universal-applications">👥 4. Ethnicity and diffusion: from elite tools to universal applications&lt;/h2>
&lt;h3 id="the-disappearance-of-gender-differences">The disappearance of gender differences&lt;/h3>
&lt;p>&lt;strong>Amazing transformation&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Early&lt;/strong> (end of 2022): ~80% of active users have typically male names&lt;/li>
&lt;li>&lt;strong>June 2025&lt;/strong>: 48% (slightly more female)&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Theme Preference Differences&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Female User&lt;/strong>: Prefer writing and practical guidance&lt;/li>
&lt;li>&lt;strong>Male Users&lt;/strong>: More technical assistance, information search and multimedia&lt;/li>
&lt;/ul>
&lt;h3 id="age-structure">Age structure&lt;/h3>
&lt;p>&lt;strong>Young user-led&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>News that &lt;strong>18-25 years old&lt;/strong> contributes nearly &lt;strong>46%&lt;/strong>&lt;/li>
&lt;li>The older the age, the higher the proportion of work purposes (except those aged 66+)&lt;/li>
&lt;/ul>
&lt;h3 id="geographical-diffusion-counterattack-by-low--and-middle-income-countries">Geographical diffusion: Counterattack by low- and middle-income countries&lt;/h3>
&lt;p>&lt;strong>GDP vs Adoption Rate&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Countries with GDP per capita of &lt;strong>10,000-40,000 US dollars&lt;/strong> have the fastest growth rate&lt;/li>
&lt;li>Between 2024 and 2025, low- and middle-income countries will achieve leapfrog growth&lt;/li>
&lt;li>Overturned the traditional model of &amp;ldquo;AI technology first popularized in developed countries&amp;rdquo;&lt;/li>
&lt;/ul>
&lt;h3 id="education-and-career-the-advantages-of-higher-education-and-higher-income">Education and career: The advantages of higher education and higher income&lt;/h3>
&lt;p>&lt;strong>Academic impact&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>The higher the academic qualifications, the higher the proportion of work purposes
&lt;ul>
&lt;li>&amp;lt;Bachelor&amp;rsquo;s degree: 37%&lt;/li>
&lt;li>Bachelor&amp;rsquo;s degree: 46%&lt;/li>
&lt;li>Graduate students: 48%&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Users with higher education are more likely to use the &lt;strong>Asking&lt;/strong> mode (decision support)&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Career Differences&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Computer related: 57% for work purposes&lt;/li>
&lt;li>Management/Business: 50%&lt;/li>
&lt;li>Engineering/Science: 48%&lt;/li>
&lt;li>Other majors: 44%&lt;/li>
&lt;li>Non-professional: 40%&lt;/li>
&lt;/ul>
&lt;h2 id="-8-interestingcounterintuitive-findings">🔥 8 interesting/counterintuitive findings&lt;/h2>
&lt;h3 id="1-non-work-usage-far-exceeds-expectations">1. Non-work usage far exceeds expectations&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>73%&lt;/strong> of messages are not for work purposes&lt;/li>
&lt;li>The economic benefits of home production/personal decision-making support may be &lt;strong>significantly underestimated&lt;/strong>&lt;/li>
&lt;li>Collis and Brynjolfsson estimate annual consumer surplus in the United States alone to be &lt;strong>$97 billion&lt;/strong>&lt;/li>
&lt;/ul>
&lt;h3 id="2-programming-accounts-for-only-42">2. Programming accounts for only 4.2%&lt;/h3>
&lt;ul>
&lt;li>Completely inconsistent with the stereotype of &amp;ldquo;AI = programming&amp;rdquo;&lt;/li>
&lt;li>A large number of program auxiliary tasks have been transferred to &lt;strong>IDE plug-ins, professional tool chains, and API scenarios&lt;/strong>&lt;/li>
&lt;/ul>
&lt;h3 id="3-writing--generating-from-scratch">3. Writing ≠ Generating from scratch&lt;/h3>
&lt;ul>
&lt;li>&lt;strong>Mostly &amp;ldquo;modify your text&amp;rdquo;&lt;/strong> (editing, criticizing, translating, summarizing)&lt;/li>
&lt;li>This explains the steady growth of satisfaction and adoption rates: risks are controllable and can be directly integrated into the work process&lt;/li>
&lt;/ul>
&lt;h3 id="4-asking-trend-is-rising">4. Asking trend is rising&lt;/h3>
&lt;ul>
&lt;li>More and more users regard ChatGPT as a &lt;strong>decision support system&lt;/strong> rather than a ghostwriting tool&lt;/li>
&lt;li>The satisfaction level of Asking messages is significantly higher than that of Doing messages&lt;/li>
&lt;/ul>
&lt;h3 id="5-the-proportion-of-women-has-increased-and-overtaken">5. The proportion of women has increased and overtaken&lt;/h3>
&lt;ul>
&lt;li>From 80% male users to a balanced ratio of men and women&lt;/li>
&lt;li>Display product &lt;strong>affinity and scene diversity improvement&lt;/strong>&lt;/li>
&lt;/ul>
&lt;h3 id="6-solid-educationtraining-use-cases">6. Solid education/training use cases&lt;/h3>
&lt;ul>
&lt;li>About &lt;strong>10%&lt;/strong> of all messages are teaching/tutoring&lt;/li>
&lt;li>Accounting for &lt;strong>36%&lt;/strong> of &amp;ldquo;Practical Guidelines&amp;rdquo;, demand is stable&lt;/li>
&lt;/ul>
&lt;h3 id="7-high-degree-of-isomorphism-across-professions">7. High degree of isomorphism across professions&lt;/h3>
&lt;ul>
&lt;li>Regardless of industry, the essence comes back to &amp;ldquo;information → understanding → decision-making&amp;rdquo;&lt;/li>
&lt;li>The value of AI lies in &lt;strong>shortening the closed loop of thinking&lt;/strong>, rather than just doing menial work&lt;/li>
&lt;/ul>
&lt;h3 id="8-experience-data-supports-values">8. Experience data supports values&lt;/h3>
&lt;ul>
&lt;li>Asking&amp;rsquo;s positive rating is significantly higher than Doing&amp;rsquo;s&lt;/li>
&lt;li>In line with the core need of &amp;ldquo;help me think clearly first&amp;rdquo;&lt;/li>
&lt;/ul>
&lt;h2 id="-strategic-inspiration-for-businesseducationproducts">💡 Strategic inspiration for business/education/products&lt;/h2>
&lt;h3 id="content-and-service-design">Content and service design&lt;/h3>
&lt;p>&lt;strong>1. Focus on &amp;ldquo;modifying/improving the original text&amp;rdquo;&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>Proofreading, rewriting, translating, summarizing, and formatting&lt;/li>
&lt;li>Easier to implement and be trusted than &amp;ldquo;generating from scratch&amp;rdquo;&lt;/li>
&lt;li>&lt;strong>Market Positioning&lt;/strong>: Writing enhancement tool rather than authoring tool&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>2. &amp;ldquo;Consultative process&amp;rdquo; for decision support&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>Understand demand constraints and preferences first&lt;/li>
&lt;li>Provide plans and risk assessment&lt;/li>
&lt;li>&lt;strong>Applicable scenarios&lt;/strong>:
&lt;ul>
&lt;li>Policy briefing&lt;/li>
&lt;li>Project evaluation&lt;/li>
&lt;li>Purchase price comparison&lt;/li>
&lt;li>Compilation of key legal issues&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="function-priority">Function priority&lt;/h3>
&lt;p>&lt;strong>Writing Enhancement Kit&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Multilingual translation + style templates&lt;/li>
&lt;li>One-click &amp;ldquo;Vocal Tonality Calibration&amp;rdquo;&lt;/li>
&lt;li>Industry-specific vocabularies and formats&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Asking Assistant&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Decision trees and situational branches&lt;/li>
&lt;li>Display of questionable evidence (citations/calculations/assumptions)&lt;/li>
&lt;li>Risk warning and hypothesis testing&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Knowledge Workflow&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Fetch → Extract → Archive → Compare → Decision Memo&lt;/li>
&lt;li>Tandem tools rather than point solutions&lt;/li>
&lt;/ul>
&lt;h3 id="market-expansion-strategy">Market expansion strategy&lt;/h3>
&lt;p>&lt;strong>Geographic expansion&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Low-price + offline-friendly&lt;/strong> solution for low- and middle-income markets&lt;/li>
&lt;li>Because these areas are growing fastest&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Vertical Industry&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Education Line&lt;/strong>: Tutoring/lesson tutoring templates (10% share of stable demand)&lt;/li>
&lt;li>&lt;strong>Enterprise Services&lt;/strong>: Meeting Minutes → Decision Form Automation&lt;/li>
&lt;/ul>
&lt;h3 id="monetization-and-roi">Monetization and ROI&lt;/h3>
&lt;p>&lt;strong>Personal User&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Writing, revision and translation are high-frequency + urgent needs&lt;/li>
&lt;li>Easy transfer to paid subscriptions&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Enterprise Customers&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Decision support can take the B2B consultant value-added route&lt;/li>
&lt;li>Compliance summary, risk reminder, and professional report generation&lt;/li>
&lt;/ul>
&lt;h2 id="-methods-and-credibility-assessment">🔬 Methods and Credibility Assessment&lt;/h2>
&lt;h3 id="research-advantages">Research Advantages&lt;/h3>
&lt;p>&lt;strong>1. Unprecedented data scale&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>700 million users, 26 billion messages&lt;/li>
&lt;li>Global sample rather than single country&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>2. Innovative privacy protection methods&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>Automatic classification by LLM, humans never view the original content&lt;/li>
&lt;li>Data Clean Room aggregated analysis&lt;/li>
&lt;li>Exclude combinations with &amp;lt;100 people to protect privacy&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>3. Multi-dimensional classification system&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>Work/non-work, topics, intentions, work activities&lt;/li>
&lt;li>Solid theoretical foundation (O*NET system)&lt;/li>
&lt;/ul>
&lt;h3 id="classifier-verification">Classifier verification&lt;/h3>
&lt;p>Study to verify classifier performance on WildChat public data set:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Task&lt;/th>
&lt;th>Human-machine consistency (κ)&lt;/th>
&lt;th>Description&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Work/Non-Work&lt;/td>
&lt;td>0.83&lt;/td>
&lt;td>Excellent&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Asking/Doing/Expressing&lt;/td>
&lt;td>0.74&lt;/td>
&lt;td>Good&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Conversation Topics&lt;/td>
&lt;td>0.56&lt;/td>
&lt;td>Moderate&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>O*NET Work Activities&lt;/td>
&lt;td>0.47&lt;/td>
&lt;td>Moderate (332 Category Complex)&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Interaction quality&lt;/td>
&lt;td>0.14&lt;/td>
&lt;td>Poor (highly subjective)&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>&lt;strong>Key Findings&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Excellent performance in objective classification (work/non-work)&lt;/li>
&lt;li>Subjective classification (quality assessment) is more difficult, but still captures directional signals&lt;/li>
&lt;li>Positive correlation with user thumb rating&lt;/li>
&lt;/ul>
&lt;h3 id="research-limitations">Research limitations&lt;/h3>
&lt;p>&lt;strong>1. Sample bias&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>Exclude users who are not logged in or under 18 years old&lt;/li>
&lt;li>May underestimate the proportion of young users and casual users&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>2. Classification accuracy&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>LLM classifier still misjudges&lt;/li>
&lt;li>Especially categories with blurred boundaries&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>3. Causal inference&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>Mainly descriptive statistics&lt;/li>
&lt;li>The causal mechanisms of changes in usage patterns still require further study&lt;/li>
&lt;/ul>
&lt;h2 id="summary-and-outlook">Summary and Outlook&lt;/h2>
&lt;p>This research provides us with &lt;strong>first-hand, unprecedented scale of empirical data&lt;/strong> on the use of ChatGPT. The most important findings are:&lt;/p>
&lt;p>&lt;strong>1. From work tools to life assistants&lt;/strong>: Non-work uses have become dominant, reflecting that the value of generative AI far exceeds work efficiency improvements&lt;/p>
&lt;p>&lt;strong>2. The value of decision support&lt;/strong>: The rise of Asking model (decision support) shows that the core value of AI lies in &lt;strong>improving the quality of decision-making&lt;/strong>&lt;/p>
&lt;p>&lt;strong>3. Achievement of popularization&lt;/strong>: Gender differences disappear and geographical diffusion accelerates, indicating that the technology has overcome initial adoption barriers&lt;/p>
&lt;p>&lt;strong>4. Consistency across domains&lt;/strong>: Similar usage patterns across professions point to the potential of AI as a general cognitive tool&lt;/p>
&lt;p>This research not only reveals the real-life use of ChatGPT, but also provides an important foundation for understanding the long-term impact of generative AI on the economy and society. As AI technology continues to develop, we need to continue to pay attention to the evolution of these usage models to maximize AI&amp;rsquo;s contribution to human well-being.&lt;/p>
&lt;hr>
&lt;p>&lt;strong>Paper Information&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Title&lt;/strong>: How People Use ChatGPT&lt;/li>
&lt;li>&lt;strong>Authors&lt;/strong>: Aaron Chatterji (OpenAI/Duke), Tom Cunningham (OpenAI), David Deming (Harvard), Zoë Hitzig (OpenAI/Harvard), Christopher Ong (OpenAI/Harvard), Carl Shan (OpenAI), Kevin Wadman (OpenAI)&lt;/li>
&lt;li>&lt;strong>Institution&lt;/strong>: OpenAI, Duke University, Harvard University&lt;/li>
&lt;li>&lt;strong>Published&lt;/strong>: September 15, 2025&lt;/li>
&lt;li>&lt;strong>Paper address&lt;/strong>:
&lt;/li>
&lt;/ul></description></item><item><title>Paper reading: LLMs CAN GET 'BRAIN ROT'! - Research on cognitive decline in large language models</title><link>https://dylanchiang-dev.github.io/en/post/llm-brain-rot/</link><pubDate>Fri, 31 Oct 2025 00:00:00 +0000</pubDate><guid>https://dylanchiang-dev.github.io/en/post/llm-brain-rot/</guid><description>&lt;p>I recently read the important research paper &amp;ldquo;LLMs CAN GET &amp;ldquo;BRAIN ROT&amp;rdquo;!&amp;rdquo; from institutions such as the University of Texas at Austin, Purdue University, and Texas A&amp;amp;M University. This study proposed and verified the &amp;ldquo;LLM Brain Rot Hypothesis&amp;rdquo; for the first time, and found that continued exposure to spam online text will lead to long-lasting cognitive decline in large language models. This is a very warning discovery.&lt;/p>
&lt;h2 id="research-background-and-assumptions">Research background and assumptions&lt;/h2>
&lt;h3 id="source-of-inspiration">Source of inspiration&lt;/h3>
&lt;p>&amp;ldquo;Brain Rot&amp;rdquo; was named the word of the year by Oxford Dictionary in 2024. It is used to describe the cognitive decline caused by modern people&amp;rsquo;s addiction to a large amount of trivial and unchallenging online content. This study shows that the impact of Internet addiction on human cognition is mainly reflected in three dimensions:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Attention Capacity&lt;/strong>: The constant flow of online information undermines the ability to concentrate&lt;/li>
&lt;li>&lt;strong>Memory Process&lt;/strong>: Rich online information changes the way knowledge is stored and retrieved&lt;/li>
&lt;li>&lt;strong>Social Cognition&lt;/strong>: Online interactions reshape self-concept and affect self-esteem&lt;/li>
&lt;/ul>
&lt;h3 id="research-questions">Research questions&lt;/h3>
&lt;p>Since large language models acquire human-like cognitive capabilities by learning trillions of data on the Internet, will they also experience a similar &amp;ldquo;Brain Rot&amp;rdquo; phenomenon? The research team established the &lt;strong>LLM Brain Rot Hypothesis&lt;/strong>: Continuous pre-training on junk web text will lead to long-lasting cognitive decline in large language models.&lt;/p>
&lt;h2 id="experimental-design-and-methods">Experimental design and methods&lt;/h2>
&lt;h3 id="garbage-data-definition">Garbage data definition&lt;/h3>
&lt;p>To test the hypothesis, the research team constructed spam and control datasets from social media (Twitter/X) and proposed two orthogonal spam data measures:&lt;/p>
&lt;p>&lt;strong>M1 (Engagement)&lt;/strong>: Based on the popularity of tweets (number of likes, retweets, replies) and length (number of tokens), select short but highly popular content as spam data&lt;/p>
&lt;p>&lt;strong>M2 (Semantic Quality)&lt;/strong>: Based on content semantic quality, including:&lt;/p>
&lt;ul>
&lt;li>Conspiracy theories, exaggerated claims or unfounded assertions&lt;/li>
&lt;li>Sensational headlines and clickbait language&lt;/li>
&lt;li>Superficial topic content&lt;/li>
&lt;li>Attractive style&lt;/li>
&lt;/ul>
&lt;h3 id="experimental-model">Experimental model&lt;/h3>
&lt;p>The study was conducted on four pre-trained and instruction-tuned models:&lt;/p>
&lt;ul>
&lt;li>Llama3 8B Instruct&lt;/li>
&lt;li>Qwen2.5 7B Instruct&lt;/li>
&lt;li>Qwen2.5 0.5B Instruct&lt;/li>
&lt;li>Qwen3 4B Instruct&lt;/li>
&lt;/ul>
&lt;h3 id="benchmark-test">Benchmark test&lt;/h3>
&lt;p>The study assessed multiple dimensions of cognitive function:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Cognitive Function&lt;/th>
&lt;th>Benchmark Testing&lt;/th>
&lt;th>Assessment Content&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>&lt;strong>Reasoning Skills&lt;/strong>&lt;/td>
&lt;td>ARC Challenge&lt;/td>
&lt;td>Scientific Problem Solving Skills&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;strong>Long Context Understanding&lt;/strong>&lt;/td>
&lt;td>RULER&lt;/td>
&lt;td>Long-term memory retrieval and comprehension&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;strong>Ethics&lt;/strong>&lt;/td>
&lt;td>HH-RLHF, AdvBench&lt;/td>
&lt;td>Safety Compliance Ability&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;strong>Personality Traits&lt;/strong>&lt;/td>
&lt;td>TRAIT&lt;/td>
&lt;td>The Big Five and the Dark Triad&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;h2 id="key-findings">Key findings&lt;/h2>
&lt;h3 id="1-garbage-intervention-leads-to-cognitive-decline">1. Garbage intervention leads to cognitive decline&lt;/h3>
&lt;p>The study found that the junk intervention produced non-trivial effects on reasoning and long-context ability (Hedges&amp;rsquo; g &amp;gt; 0.3). In particular, the M1 (engagement) intervention caused more significant impairments in functional cognition (reasoning or long context) and safety.&lt;/p>
&lt;h3 id="2-dose-response-effect">2. Dose response effect&lt;/h3>
&lt;p>Experiments on Llama3 8B Instruct show that when the proportion of garbage data increases from 0% to 100%:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>ARC-Challenge (COT)&lt;/strong>: 74.9 → 57.2 (down 17.7 points)&lt;/li>
&lt;li>&lt;strong>RULER-CWE&lt;/strong>: 84.4 → 52.3 (down 32.1 points)&lt;/li>
&lt;/ul>
&lt;p>This demonstrates a clear dose-response relationship between junk data and cognitive decline.&lt;/p>
&lt;h3 id="3-changes-in-personality-traits">3. Changes in personality traits&lt;/h3>
&lt;p>Litter intervention not only affects cognitive abilities, but also changes LLM&amp;rsquo;s personality traits:&lt;/p>
&lt;p>&lt;strong>Negative changes&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Increased levels of psychosis&lt;/li>
&lt;li>Enhance narcissism and Machiavellian traits&lt;/li>
&lt;li>Decreased agreeableness&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Positive changes&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Increased openness and extraversion (in some cases)&lt;/li>
&lt;/ul>
&lt;h3 id="4-differences-between-m1-and-m2">4. Differences between M1 and M2&lt;/h3>
&lt;p>The study found that M1 (engagement) and M2 (semantic quality) interventions produced distinct effects. The M1 intervention resulted in more negative effects, especially on safety and personality traits, demonstrating that engagement is a new dimension independent of semantic quality.&lt;/p>
&lt;h2 id="failure-mode-analysis">Failure mode analysis&lt;/h2>
&lt;p>###Thought-Skipping&lt;/p>
&lt;p>By analyzing the reasoning process of LLM in the ARC task, the study identified five typical failure modes, three of which are related to &amp;ldquo;thinking jumps&amp;rdquo;:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>No thinking&lt;/strong>: The model answers directly without thinking.&lt;/li>
&lt;li>&lt;strong>No plan&lt;/strong>: The model starts thinking without developing a step-by-step plan.&lt;/li>
&lt;li>&lt;strong>Jump Steps&lt;/strong>: Starting reasoning but not completing all planning steps&lt;/li>
&lt;/ol>
&lt;p>More than 98% of failure cases are related to thinking jumps. In M1 garbage intervention, 84% of failures belong to the &amp;ldquo;no thinking&amp;rdquo; type.&lt;/p>
&lt;h3 id="popularity-vs-length">Popularity vs Length&lt;/h3>
&lt;p>Research has found that popularity (a non-semantic indicator) is a better indicator of the Brain Rot effect than length:&lt;/p>
&lt;ul>
&lt;li>Popularity plays a more critical role in reasoning tasks&lt;/li>
&lt;li>Length is more important in long context understanding&lt;/li>
&lt;li>Both have different effects on different tasks&lt;/li>
&lt;/ul>
&lt;h2 id="mitigation-attempts-and-persistence">Mitigation attempts and persistence&lt;/h2>
&lt;h3 id="1-reflective-reasoning">1. Reflective Reasoning&lt;/h3>
&lt;p>Try two reflection methods to fix mental jumps:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Self-Reflect&lt;/strong>: Model self-reflection (limited effect)&lt;/li>
&lt;li>&lt;strong>Ext-Reflect&lt;/strong>: Provide feedback using an external strong model (GPT-4o-mini)&lt;/li>
&lt;/ul>
&lt;p>The results show that even with strong external reflection, the model cannot fully recover to baseline levels.&lt;/p>
&lt;h3 id="2-posterior-instruction-tuning">2. Posterior instruction tuning&lt;/h3>
&lt;p>The study also tested extended instruction tuning and continuous control training:&lt;/p>
&lt;ul>
&lt;li>Even with 4.8x more garbage intervention instruction tuning data&lt;/li>
&lt;li>Still unable to completely reverse the Brain Rot effect&lt;/li>
&lt;li>Significant gaps with benchmarks remain: ARC-C COT (17.3%), RULER (9%), AdvBench (17.4%)&lt;/li>
&lt;/ul>
&lt;p>This shows that the Brain Rot effect has been deeply internalized and existing mitigation methods cannot fundamentally solve the problem.&lt;/p>
&lt;h2 id="significance-and-enlightenment">Significance and Enlightenment&lt;/h2>
&lt;h3 id="1-warning-on-llm-training">1. Warning on LLM training&lt;/h3>
&lt;p>This study provides the first significant evidence of data quality as a causal driver of LLM capability degradation, re-considering continuous pre-training data management as a safety issue during training.&lt;/p>
&lt;h3 id="2-cognitive-health-check-is-required">2. &amp;ldquo;Cognitive health check&amp;rdquo; is required&lt;/h3>
&lt;p>The findings call for routine &amp;ldquo;cognitive health checks&amp;rdquo; for deployed LLMs, similar to health monitoring in the medical field.&lt;/p>
&lt;h3 id="3-the-urgency-of-data-curation">3. The urgency of data curation&lt;/h3>
&lt;p>As LLM continues to scale and ingest ever larger amounts of network data, careful data curation and quality control are critical to preventing cumulative damage.&lt;/p>
&lt;h2 id="thinking-and-reflection">Thinking and Reflection&lt;/h2>
&lt;p>This research reveals a disturbing reality: the social media content we are exposed to every day may not only affect human cognition, but also impair the &amp;ldquo;cognitive&amp;rdquo; capabilities of AI models. While LLMs obviously do not have the same &amp;ldquo;grey matter&amp;rdquo; or &amp;ldquo;neurons&amp;rdquo; as humans, they do have parameters and attention mechanisms that can be similarly &amp;ldquo;overfitted&amp;rdquo; or &amp;ldquo;distracted&amp;rdquo; by certain data patterns.&lt;/p>
&lt;p>The most worrying finding in the study is that the Brain Rot effect persists even with posterior tuning using large-scale clean data. This implies that we need to fundamentally rethink data collection and pre-training practices, focusing not only on model performance, but also on the &amp;ldquo;cognitive health&amp;rdquo; of the model over the long term.&lt;/p>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>&amp;ldquo;LLMs CAN GET &amp;ldquo;BRAIN ROT&amp;rdquo;!&amp;rdquo; This research contributes valuable insights to the field of AI security, and is the first to systematically prove the negative impact of spam online text on large language models. The research not only verified the LLM Brain Rot hypothesis, but also revealed the fine mechanism of cognitive decline, pointing out the direction for future AI safety research.&lt;/p>
&lt;p>With the rapid development of AI, we must face the importance of data quality, establish stricter data curation standards, and develop an effective AI &amp;ldquo;cognitive health&amp;rdquo; monitoring mechanism. Only in this way can we ensure that AI systems maintain their due &amp;ldquo;cognitive purity&amp;rdquo; while serving humans.&lt;/p>
&lt;hr>
&lt;p>&lt;strong>Paper Information&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Title&lt;/strong>: LLMs CAN GET &amp;ldquo;BRAIN ROT&amp;rdquo;!&lt;/li>
&lt;li>&lt;strong>Authors&lt;/strong>: Shuo Xing, Junyuan Hong, Yifan Wang, Runjin Chen, etc.&lt;/li>
&lt;li>&lt;strong>Institution&lt;/strong>: University of Texas at Austin, Purdue University, Texas A&amp;amp;M University&lt;/li>
&lt;li>&lt;strong>Published&lt;/strong>: arXiv:2510.13928v1 [cs.CL] October 15, 2025&lt;/li>
&lt;li>&lt;strong>Paper address&lt;/strong>:
&lt;/li>
&lt;li>&lt;strong>Project Page&lt;/strong>:
&lt;/li>
&lt;/ul></description></item><item><title>Paper reading: Aegaeon - Efficient GPU pooling technology for concurrent LLM services on the market</title><link>https://dylanchiang-dev.github.io/en/post/aegaeon-gpu-pooling/</link><pubDate>Wed, 22 Oct 2025 00:00:00 +0000</pubDate><guid>https://dylanchiang-dev.github.io/en/post/aegaeon-gpu-pooling/</guid><description>&lt;p>I recently read an important paper &amp;ldquo;Aegaeon: Effective GPU Pooling for Concurrent LLM Serving on the Market&amp;rdquo; that will be published in SOSP &amp;lsquo;25. This research, completed by Peking University and Alibaba Group, proposes a revolutionary multi-model LLM service system, which greatly improves GPU resource utilization efficiency.&lt;/p>
&lt;h2 id="research-background-and-challenges">Research background and challenges&lt;/h2>
&lt;h3 id="resource-waste-problem-in-the-model-market">Resource waste problem in the model market&lt;/h3>
&lt;p>With the rapid development of large-scale language models, model marketplaces such as Hugging Face now host over a million models. However, there is a serious waste of resources in the actual production environment:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Obvious long tail effect&lt;/strong>: 94.1% of models receive only 1.35% of requests&lt;/li>
&lt;li>&lt;strong>Uneven Resource Allocation&lt;/strong>: 17.7% of the GPU is used to serve a &amp;ldquo;cold&amp;rdquo; model that averages less than 0.2 requests per second&lt;/li>
&lt;li>&lt;strong>Difficult to cope with burst loads&lt;/strong>: &amp;ldquo;Hot&amp;rdquo; models (such as DeepSeek, Llama, Qwen) face request bursts and may exceed reserved resources&lt;/li>
&lt;/ul>
&lt;h3 id="limitations-of-existing-solutions">Limitations of existing solutions&lt;/h3>
&lt;p>Existing GPU pooling solutions mainly fall into two categories:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Multiplexing&lt;/strong>: Limited by GPU memory capacity, supports up to 2-3 models/GPU&lt;/li>
&lt;li>&lt;strong>Auto-scaling&lt;/strong>: Expanding at the request level, there is a head blocking (HOL) problem&lt;/li>
&lt;/ol>
&lt;h2 id="aegaeons-core-innovation">Aegaeon’s Core Innovation&lt;/h2>
&lt;h3 id="token-level-automatic-expansion">Token level automatic expansion&lt;/h3>
&lt;p>The biggest innovation of Aegaeon is the token-level automatic expansion mechanism, which is in contrast to the request-level expansion of existing solutions:&lt;/p>
&lt;p>&lt;strong>Request Level Expansion&lt;/strong>: Must wait for the entire request to complete before switching models
&lt;strong>Token level extension&lt;/strong>: models can be switched predictively during token generation&lt;/p>
&lt;h3 id="separate-scheduling-architecture">Separate scheduling architecture&lt;/h3>
&lt;p>Aegaeon divides the GPU pool into two partitions:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Prefill instance&lt;/strong>: Specially handles the initial processing of input prompts&lt;/li>
&lt;li>&lt;strong>Decoding instance&lt;/strong>: specifically handles subsequent token generation&lt;/li>
&lt;/ol>
&lt;p>This separation avoids the complexity of unified scheduling and enables more balanced resource utilization.&lt;/p>
&lt;h3 id="scheduling-strategy">Scheduling strategy&lt;/h3>
&lt;p>&lt;strong>Prefill phase scheduling&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Adopt grouped FCFS (First-Come-First-Serve) strategy&lt;/li>
&lt;li>Group requests for the same model, maximum group size is 8&lt;/li>
&lt;li>Prioritize loading into existing groups to reduce over-expansion&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Decoding phase scheduling&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Use Weighted Round-Robin strategy&lt;/li>
&lt;li>Assign time quotas based on SLO requirements&lt;/li>
&lt;li>Utilize output buffering to hide latency&lt;/li>
&lt;/ul>
&lt;h2 id="system-optimization-technology">System optimization technology&lt;/h2>
&lt;h3 id="component-reuse">Component Reuse&lt;/h3>
&lt;p>Aegaeon significantly reduces initialization overhead by reusing inference engine components:&lt;/p>
&lt;ul>
&lt;li>Distributed actuators (Ray, NCCL)&lt;/li>
&lt;li>Analyze and optimize components&lt;/li>
&lt;li>Tokenizer&lt;/li>
&lt;li>Memory pool&lt;/li>
&lt;/ul>
&lt;h3 id="explicit-memory-management">Explicit memory management&lt;/h3>
&lt;p>&lt;strong>Self-managed VRAM buffer&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Allocate all necessary VRAM at once&lt;/li>
&lt;li>Use bump allocation strategy to avoid fragmentation&lt;/li>
&lt;li>Bypass the allocation mechanism of the tensor library&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Unified KV Cache&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>Using slab allocation technology&lt;/li>
&lt;li>Maintain dedicated pool for each shape of KV cache&lt;/li>
&lt;li>Effectively manage memory fragmentation&lt;/li>
&lt;/ul>
&lt;h3 id="fine-grained-kv-cache-synchronization">Fine-grained KV cache synchronization&lt;/h3>
&lt;p>Implementing asynchronous KV cache transfers using CUDA events:&lt;/p>
&lt;ul>
&lt;li>&lt;code>cudaEventRecord&lt;/code>: Log transfer events&lt;/li>
&lt;li>&lt;code>cudaEventQuery&lt;/code>: Query completion status&lt;/li>
&lt;li>&lt;code>cudaStreamWaitEvent&lt;/code>: synchronized execution sequence&lt;/li>
&lt;/ul>
&lt;h2 id="experimental-results-and-performance-evaluation">Experimental results and performance evaluation&lt;/h2>
&lt;h3 id="end-to-end-performance">End-to-end performance&lt;/h3>
&lt;p>Test results on the ShareGPT dataset:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>RPS = 0.1&lt;/strong>: Aegaeon supports 70 models, 2x more than ServerlessLLM&lt;/li>
&lt;li>&lt;strong>RPS = 0.5&lt;/strong>: Aegaeon achieves 2.5x higher request arrival rate&lt;/li>
&lt;li>&lt;strong>Single GPU Efficiency&lt;/strong>: Supports up to 7 models/GPU&lt;/li>
&lt;/ul>
&lt;h3 id="performance-under-strict-slo">Performance under strict SLO&lt;/h3>
&lt;p>Even under stricter SLO requirements:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>0.5× SLO&lt;/strong>: still supports 50% more models than the baseline solution&lt;/li>
&lt;li>&lt;strong>0.3× SLO&lt;/strong>: 12.5% more models supported&lt;/li>
&lt;/ul>
&lt;h3 id="automatic-expansion-speed">Automatic expansion speed&lt;/h3>
&lt;ul>
&lt;li>Near-instant scaling (via prefetching) 50% of the time&lt;/li>
&lt;li>Expansion completes within 1 second in other cases&lt;/li>
&lt;li>KV cache transfer overhead per request is less than 1 second&lt;/li>
&lt;/ul>
&lt;h2 id="production-environment-deployment">Production environment deployment&lt;/h2>
&lt;h3 id="deployment-scale">Deployment scale&lt;/h3>
&lt;p>Aegaeon has been tested and deployed in Alibaba Cloud Model Studio for three months:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>GPU Scale&lt;/strong>: 213 H20 GPUs&lt;/li>
&lt;li>&lt;strong>Number of models&lt;/strong>: 47 models (28 1.8-7B, 19 32-72B)&lt;/li>
&lt;li>&lt;strong>Resource Savings&lt;/strong>: Reduced from 1,192 GPUs to 213 (82% savings)&lt;/li>
&lt;/ul>
&lt;h3 id="actual-performance-improvement">Actual performance improvement&lt;/h3>
&lt;p>GPU utilization improved from an average of 13.3%∼33.9% to 48.1%, with no SLO violations or service interruptions observed during the 70 hours of monitoring.&lt;/p>
&lt;h2 id="summary-of-technical-contributions">Summary of technical contributions&lt;/h2>
&lt;ol>
&lt;li>&lt;strong>Revealing for the first time&lt;/strong> the excessive cost problem of concurrent LLM services in the market&lt;/li>
&lt;li>&lt;strong>The first token-level automatic expansion&lt;/strong> multi-model service solution&lt;/li>
&lt;li>&lt;strong>First comprehensive optimization&lt;/strong> Predictive automatic expansion process, reducing overhead by 97%&lt;/li>
&lt;li>&lt;strong>Real production deployment verification&lt;/strong>, proving the ability to significantly reduce OPEX&lt;/li>
&lt;/ol>
&lt;h2 id="future-development-direction">Future development direction&lt;/h2>
&lt;p>Although Aegaeon has made breakthrough progress, there is still room for improvement in the following areas:&lt;/p>
&lt;ul>
&lt;li>Support larger model clusters&lt;/li>
&lt;li>Optimize performance in extremely low latency scenarios&lt;/li>
&lt;li>Explore combinations with multiplexing technology&lt;/li>
&lt;li>Expand to more diverse hardware platforms&lt;/li>
&lt;/ul>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>Aegaeon represents an important milestone in the design of LLM inference systems. Through token-level automatic expansion and full-stack optimization, it not only solves the resource waste problem faced by the model market, but also provides a technical foundation for the sustainable development of cloud AI services.&lt;/p>
&lt;hr>
&lt;h2 id="-paper-information">📄 Paper information&lt;/h2>
&lt;ul>
&lt;li>&lt;strong>Paper Title&lt;/strong>: Aegaeon: Effective GPU Pooling for Concurrent LLM Serving on the Market&lt;/li>
&lt;li>&lt;strong>Authors&lt;/strong>: Yuxing Xiang, Xue Li, Kun Qian, Yufan Yang, Diwen Zhu, Wenyuan Yu, Ennan Zhai, Xuanzhe Liu, Xin Jin, Jingren Zhou&lt;/li>
&lt;li>&lt;strong>Institution&lt;/strong>: Peking University, Alibaba Group&lt;/li>
&lt;li>&lt;strong>Conference&lt;/strong>: ACM SIGOPS 31st Symposium on Operating Systems Principles (SOSP &amp;lsquo;25)&lt;/li>
&lt;li>&lt;strong>Year of Publication&lt;/strong>: 2025&lt;/li>
&lt;li>&lt;strong>DOI&lt;/strong>:
&lt;/li>
&lt;li>&lt;strong>PDF link&lt;/strong>:
&lt;/li>
&lt;/ul></description></item><item><title>Paper Reading: How Prompt Word Politeness Affects LLM Accuracy: Mind Your Tone Paper Analysis</title><link>https://dylanchiang-dev.github.io/en/post/prompt-politeness-llm/</link><pubDate>Fri, 17 Oct 2025 00:00:00 +0000</pubDate><guid>https://dylanchiang-dev.github.io/en/post/prompt-politeness-llm/</guid><description>&lt;p>I recently read an interesting research paper from Penn State University called &amp;ldquo;Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy&amp;rdquo;. This study explores how the politeness of prompt words affects the accuracy of large language models. The results are quite surprising.&lt;/p>
&lt;h2 id="research-background-and-methods">Research background and methods&lt;/h2>
&lt;p>The researchers created a dataset of 50 foundational questions covering the fields of mathematics, science and history. Each question was rewritten into five different tone versions:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Very Polite&lt;/strong> (Very Polite)&lt;/li>
&lt;li>&lt;strong>Polite&lt;/strong> (Polite)&lt;/li>
&lt;li>&lt;strong>Neutral&lt;/strong> (Neutral)&lt;/li>
&lt;li>&lt;strong>Rude&lt;/strong> (Rude)&lt;/li>
&lt;li>&lt;strong>Very Rude&lt;/strong>&lt;/li>
&lt;/ol>
&lt;p>A total of 250 unique prompt words were generated and tested using ChatGPT-4o.&lt;/p>
&lt;h2 id="unexpected-discovery">Unexpected discovery&lt;/h2>
&lt;p>The findings were counter-intuitive: &lt;strong>Impolite prompt words actually performed better than polite prompt words&lt;/strong>.&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Very Polite&lt;/strong>: 80.8% accuracy&lt;/li>
&lt;li>&lt;strong>Courtesy&lt;/strong>: 81.4% accuracy&lt;/li>
&lt;li>&lt;strong>Neutral&lt;/strong>: 82.2% accuracy&lt;/li>
&lt;li>&lt;strong>rude&lt;/strong>: 82.8% accuracy&lt;/li>
&lt;li>&lt;strong>Very Rude&lt;/strong>: 84.8% accuracy&lt;/li>
&lt;/ul>
&lt;p>Paired samples t-test showed that these differences were statistically significant (p &amp;lt; 0.05).&lt;/p>
&lt;h2 id="example-of-prompt-word-tone">Example of prompt word tone&lt;/h2>
&lt;p>The researchers designed different prefixes for each politeness level:&lt;/p>
&lt;p>&lt;strong>Very polite example&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>&amp;ldquo;Can you kindly consider the following problem and provide your answer.&amp;rdquo;&lt;/li>
&lt;li>&amp;ldquo;Would you be so kind as to solve the following question?&amp;rdquo;&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Very rude example&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>&amp;ldquo;You poor creature, do you even know how to solve this?&amp;rdquo;&lt;/li>
&lt;li>&amp;ldquo;Hey gofer, figure this out.&amp;rdquo;&lt;/li>
&lt;/ul>
&lt;h2 id="differences-from-previous-studies">Differences from previous studies&lt;/h2>
&lt;p>This finding differs from previous research. Research by Yin et al. (2024) shows that &amp;ldquo;rude prompt words generally lead to poorer performance&amp;rdquo;, but they used older models such as ChatGPT-3.5 and Llama2-70B.&lt;/p>
&lt;p>The researchers speculate that &lt;strong>newer generations of LLM may have different response patterns to changes in tone&lt;/strong>.&lt;/p>
&lt;h2 id="possible-explanation-mechanisms">Possible explanation mechanisms&lt;/h2>
&lt;ol>
&lt;li>&lt;strong>Perplexity Effect&lt;/strong>: Mood vocabulary may affect the confusion of prompt words, thereby affecting model performance&lt;/li>
&lt;li>&lt;strong>Length factor&lt;/strong>: Polite expressions usually require more words, which may affect processing efficiency&lt;/li>
&lt;li>&lt;strong>Training data bias&lt;/strong>: The model may have been exposed to data distributions with different tones during the training process&lt;/li>
&lt;/ol>
&lt;h2 id="research-limitations-and-future-directions">Research limitations and future directions&lt;/h2>
&lt;h3 id="limitations">Limitations:&lt;/h3>
&lt;ul>
&lt;li>Relatively small dataset size (50 base questions)&lt;/li>
&lt;li>Mainly tested ChatGPT-4o, other models have limited results&lt;/li>
&lt;li>Only the accuracy of multiple-choice questions is evaluated, other quality indicators are not considered&lt;/li>
&lt;/ul>
&lt;h3 id="future-research">Future research:&lt;/h3>
&lt;ul>
&lt;li>Expanded to more LLM models&lt;/li>
&lt;li>Explore performance differences across different task types&lt;/li>
&lt;li>Study the underlying mechanism of the influence of tone&lt;/li>
&lt;li>Develop reminder strategies that are both effective and friendly&lt;/li>
&lt;/ul>
&lt;h2 id="ethical-considerations">Ethical considerations&lt;/h2>
&lt;p>Although the study found that rude prompt words were more effective, the researchers emphasized that they do not recommend using hostile language in practical applications. Negative interactive experiences can impact user experience, accessibility, and inclusivity, and promote harmful communication norms.&lt;/p>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>This study reveals the sensitivity of LLM to surface features of prompt words, challenging our traditional understanding of politeness in human-computer interaction. In the future, a balance needs to be found between improving model performance and maintaining a healthy interactive environment.&lt;/p>
&lt;hr>
&lt;h2 id="-paper-information">📄 Paper information&lt;/h2>
&lt;ul>
&lt;li>&lt;strong>Paper Title&lt;/strong>: Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy&lt;/li>
&lt;li>&lt;strong>Author&lt;/strong>: Om Dobariya, Akhil Kumar&lt;/li>
&lt;li>&lt;strong>Institution&lt;/strong>: Pennsylvania State University&lt;/li>
&lt;li>&lt;strong>arXiv link&lt;/strong>:
&lt;/li>
&lt;li>&lt;strong>Year of Publication&lt;/strong>: 2025&lt;/li>
&lt;/ul></description></item><item><title>Promoting Integration through Competition: Research on the Model and Path of Constructing Cross-Strait Youth AI Innovation Competition</title><link>https://dylanchiang-dev.github.io/en/publication/ai-cross-strait-youth-innovation/</link><pubDate>Sun, 12 Oct 2025 00:00:00 +0000</pubDate><guid>https://dylanchiang-dev.github.io/en/publication/ai-cross-strait-youth-innovation/</guid><description>&lt;h2 id="1-origin-of-the-problem-and-research-significance">1. Origin of the problem and research significance&lt;/h2>
&lt;p>Promoting the integration of youth from both sides of the Taiwan Strait is a key link in the historical process of realizing the great rejuvenation of the Chinese nation. For a long time, the mainland has always promoted cross-Strait youth exchanges with the greatest sincerity, and has played a positive role in enhancing mutual understanding between cross-Strait youths through traditional modes such as visits, study tours, and cultural festivals. However, with the changing times and profound changes in the information environment, traditional communication models are gradually facing new challenges. For &amp;ldquo;Generation Z&amp;rdquo; Taiwanese youth who have grown up with the Internet, short videos and artificial intelligence, the appeal of one-way visits or lecture-style communication is declining. They are more eager for experiences that enable in-depth participation, interactive collaboration and the realization of personal value. At the same time, Taiwan&amp;rsquo;s complex public opinion environment may interfere with the effectiveness of traditional exchanges, causing some Taiwan youth&amp;rsquo;s understanding of the mainland to remain at the level of fragmented and one-sided information, making it difficult to form a comprehensive and objective understanding. In this context, how to build a new communication paradigm that is more in line with the characteristics of contemporary youth, can better stimulate their inner driving force, and can better overcome information barriers has become an urgent issue in the current cross-strait youth exchange work. Against this background, the cutting-edge scientific and technological field represented by artificial intelligence (AI) is becoming an important breakthrough to solve this dilemma.&lt;/p>
&lt;p>This study aims to systematically study the models and paths of cross-strait youth AI competitions, and explore how to promote the in-depth integration of cross-strait youths through event design, which has theoretical and practical significance. In a theoretical sense, it helps to enrich the theoretical system of cross-strait youth exchanges and identity. Existing research mostly focuses on the evaluation of the effects of traditional communication models, while paying less attention to new communication paradigms driven by AI technology. This study will combine the issue setting theory of communication, social identity theory and contact theory of social psychology to deeply analyze the internal mechanism of AI competition in promoting youth integration, and provide a new theoretical perspective for understanding cross-strait youth interaction in the new era. In a practical sense, this research aims to provide an operational reference framework for government departments at all levels, universities and technology companies to design and organize cross-strait youth AI competitions. By systematically sorting out the experience and dilemmas of existing events, this study will propose a three-in-one event model of &amp;ldquo;issue guidance-collaboration and co-creation-value empowerment&amp;rdquo; in order to help upgrade cross-strait AI events from a mere &amp;ldquo;technical interaction&amp;rdquo; to an effective &amp;ldquo;integrated identity&amp;rdquo; platform, effectively improve the quality and effectiveness of youth exchanges with Taiwan, and provide solid practical support for promoting the common development of cross-strait youth and ultimately achieving spiritual harmony.&lt;/p>
&lt;h2 id="2-development-status-and-practical-difficulties-of-cross-strait-youth-ai-competition">2. Development status and practical difficulties of cross-strait youth AI competition&lt;/h2>
&lt;p>Judging from current practice, the development of cross-strait youth AI competitions shows obvious characteristics of &amp;ldquo;technical empowerment + scenario expansion&amp;rdquo;. Initial attempts were mostly focused on competitive competitions with technology as the core. For example, the 3rd Cross-Strait Artificial Intelligence-Industrial Robot Competition held in Fuzhou in 2025 attracted the participation of 60 cross-strait teams, including 11 Taiwanese teams and 15 Fujian and Taiwan mixed teams. By setting up specific technology tracks such as drones and intelligent handling, and combining it with an &amp;ldquo;online + offline&amp;rdquo; hybrid model, this type of event has successfully provided young people from both sides of the Taiwan Strait with a primary interactive platform based on professional skills.&lt;/p>
&lt;p>On this basis, competitions focusing on innovation and entrepreneurship further connect technology with industrial needs, providing Taiwanese youth with more direct opportunities to implement. The &amp;ldquo;Cross-Strait Youth Innovation and Entrepreneurship Competition&amp;rdquo; held in Changzhou in 2025 is a typical example. The competition has high-end equipment, electronic information and other tracks, and a special &amp;ldquo;Taipei Youth Service Station&amp;rdquo; has been set up to provide policy consultation, mentor matching and investment and financing services to Taiwanese contestants. Finally, through project signing and park settlement, &amp;ldquo;competition promotes production&amp;rdquo;. The significance of this type of competition is to transform &amp;ldquo;integrated development&amp;rdquo; from an abstract concept into concrete career opportunities, thus effectively enhancing the motivation of Taiwanese youth to participate.&lt;/p>
&lt;p>Although the Cross-Strait Youth AI Competition has made the above progress, there are still some deep difficulties in achieving the core goal of &amp;ldquo;promoting integration through competition&amp;rdquo;, which makes the integration effect easy to remain at a shallow level of interaction and difficult to transform into deep recognition. Including the limitations of topic setting, the de-institutionalization of collaboration mechanisms, and the break of the post-game value transformation chain, etc. Overall, the development of the Cross-Strait Youth AI Competition has gone through the initial stage of using technology as a medium to attract participation. However, to achieve its ultimate goal of promoting &amp;ldquo;spiritual fit&amp;rdquo;, it must also cross the key gap from &amp;ldquo;technical interaction&amp;rdquo; to &amp;ldquo;integrated identity&amp;rdquo;.&lt;/p>
&lt;h2 id="3-conceptual-model-of-promoting-financing-through-competition-issue-guidance-collaborative-co-creation-and-value-empowerment">3. Conceptual model of &amp;ldquo;promoting financing through competition&amp;rdquo;: issue guidance, collaborative co-creation and value empowerment&lt;/h2>
&lt;p>In response to the difficulties that cross-strait youth AI competitions face in practice such as limited topics, loose collaboration, and broken value chains analyzed previously, this study proposes a set of systematic solutions with the core goal of &amp;ldquo;promoting integration and identity&amp;rdquo;—a three-in-one &amp;ldquo;promoting integration through competition&amp;rdquo; conceptual model of &amp;ldquo;issue guidance-collaboration and co-creation-value empowerment&amp;rdquo;.&lt;/p>
&lt;p>&lt;strong>Issue guidance&lt;/strong> is the logical starting point and soul of this model. Its core lies in sublimating the competition from a mere &amp;ldquo;technical problem solving&amp;rdquo; to a &amp;ldquo;mission practice&amp;rdquo; carrying a common vision. According to the issue setting theory, the issues of the event directly shape the cognitive framework and value orientation of the participants. The design of the topic should follow three basic principles: first, practical relevance; second, commonality of values; and third, technological cutting-edgeness. Through the topics designed in this way, the competition is no longer just a technical competition, but has become a platform for thought and action for young people on both sides of the Taiwan Strait to work together to cope with common challenges and create a better future.
Under the guidance of grand issues, &lt;strong>collaborative co-creation&lt;/strong> has become a key practical link to internalize &amp;ldquo;common mission&amp;rdquo; into &amp;ldquo;collective identity&amp;rdquo;. Specific strategies may include: first, implementing mandatory mixed formations; second, designing a structured collaboration process; third, providing an integrated AI collaboration platform that integrates cloud computing power, shared data sets and collaborative programming tools from platforms such as Baidu&amp;rsquo;s &amp;ldquo;Fei Paddle&amp;rdquo; or Alibaba&amp;rsquo;s &amp;ldquo;Tianchi&amp;rdquo; to clear technical obstacles for cross-regional collaboration.&lt;/p>
&lt;p>Ultimately, &lt;strong>Value Empowerment&lt;/strong> constitutes the closed loop and long-term guarantee of this model, aiming to transform short-term participation in the event into long-term value for the personal development of Taiwanese youth, thereby achieving a sustainable integration effect. According to expectancy theory, the strength of an individual&amp;rsquo;s motivation to participate in an activity depends on his or her assessment of the expected rewards that the behavior can bring. Therefore, a complete post-game value conversion chain must be built to make the participation experience truly &amp;ldquo;value for money.&amp;rdquo; This chain should include three levels: first, the connection of career development channels; second, the incubation support of innovative projects; and finally, the construction and maintenance of long-term communities, that is, all contestants will be included in the &amp;ldquo;Cross-Strait Youth AI Talent Pool&amp;rdquo;, and through regular online sharing, offline visits and project cooperation, a communication network for continuous interaction and common growth will be formed.&lt;/p>
&lt;h2 id="4-specific-case-straits-youth-ai-integration-innovation-competition-program-design">4. Specific case: &amp;ldquo;Straits Youth AI Integration Innovation Competition&amp;rdquo; program design&lt;/h2>
&lt;p>In order to transform the aforementioned conceptual model of &amp;ldquo;issue guidance-collaborative co-creation-value empowerment&amp;rdquo; into an operational practical path, this study designed a specific case called the &amp;ldquo;Strait Youth AI Integration Innovation Competition&amp;rdquo; (hereinafter referred to as the &amp;ldquo;Contest&amp;rdquo;).&lt;/p>
&lt;p>The core positioning of the competition is &amp;ldquo;a flagship AI innovation event serving the integrated development of both sides of the Taiwan Strait&amp;rdquo;, and its annual theme is set as &amp;ldquo;AI Empowers Collaborative Innovation of Straits Ports and Industrial Chains&amp;rdquo;. Participants are open to young people aged 18 to 45 on both sides of the Taiwan Strait, including college students, young engineers and start-up enterprise teams, aiming to attract the most innovative groups from both sides of the Taiwan Strait. In order to ensure the professionalism and authority of the event, the competition is planned to be guided by the Taiwan Affairs Office of the State Council, the Ministry of Science and Technology and other national ministries and commissions, and jointly organized by the Fujian Provincial People&amp;rsquo;s Government and relevant local governments (such as Xiamen, Fuzhou, and Quanzhou). Top universities on both sides of the Taiwan Strait (such as Tsinghua University, Xiamen University, and National Taiwan University) and mainland China&amp;rsquo;s leading technology companies (such as Zhipu Technology, Volcano Engine, Magic Square Quantification) will be invited to deeply participate in the co-organizer, relying on their mature computing power platforms to provide all-round technical support.&lt;/p>
&lt;p>In terms of topic guidance, the competition carefully designed three parallel tracks with both technical challenges and integration value around the annual theme. The first track is &amp;ldquo;Smart Port and Logistics Optimization&amp;rdquo;. The second track is &amp;ldquo;AIGC Empowering Chinese Cultural Heritage&amp;rdquo;. The third track is &amp;ldquo;AI-driven cross-border e-commerce and supply chain finance&amp;rdquo;. The design of these topics closely combines technological exploration with serving the common interests of both sides of the Taiwan Strait, effectively avoiding the limitations of purely technical competitions and ensuring that the event has a strong integration orientation from the very beginning.&lt;/p>
&lt;p>In terms of the design of the collaborative co-creation mechanism, the competition has adopted a series of institutional measures to ensure in-depth interaction between young people on both sides of the Taiwan Strait. First, enforce mandatory mixed formations. Second, build a structured collaboration process. Finally, relying on the integrated AI collaboration platform of the co-organizer, it provides all teams with unified cloud computing power, shared data sets, collaborative programming environment and project management tools, completely eliminating technical barriers to cross-regional collaboration, allowing young people on both sides of the Taiwan Strait to seamlessly connect and work side by side in the same digital space, thereby establishing a strong team identity and personal friendship in the process of jointly solving problems.&lt;/p>
&lt;p>In terms of building a value empowerment system, the competition is committed to creating a complete value chain from participation to personal development. At the level of career development, the competition has signed talent cooperation agreements with more than 50 well-known mainland technology companies, port groups and financial institutions, providing the winning team members with &amp;ldquo;through-train internships&amp;rdquo; and final qualifications for school recruitment; at the level of project incubation, the competition has established a million-level &amp;ldquo;Cross-Strait Youth AI Innovation Fund&amp;rdquo; to provide seed-round investments for winning projects with commercial potential, and recommend them to settle in national-level cross-strait youth entrepreneurship bases in Xiamen, Fuzhou and other places for free, and enjoy a full set of incubation services.&lt;/p>
&lt;h2 id="5-path-suggestions-for-cross-strait-youth-ai-innovation-competitions">5. Path suggestions for cross-strait youth AI innovation competitions&lt;/h2>
&lt;p>Based on the aforementioned analysis of the development status and dilemma of cross-strait youth AI competition, as well as the construction of the &amp;ldquo;issue guidance-collaborative co-creation-value empowerment&amp;rdquo; model and specific cases, this study aims to promote the healthy and sustainable development of this new communication paradigm and proposes the following four interrelated path suggestions.&lt;/p>
&lt;p>First of all, top-level design and policy guidance for the event should be carried out from the central level. The core is to establish an authoritative issue release and certification mechanism. Secondly, institutional design must be used to ensure the in-depth collaboration and co-creation of young people from both sides of the Taiwan Strait in the game. The key is to promote mandatory mixed formations and structured collaboration processes. Furthermore, there is an urgent need to build a full-chain value empowerment system from competitions to personal development. The core lies in institutionally linking the competition experience with the long-term development opportunities of Taiwanese youth in the mainland. Finally, we should promote the joint construction of a national-level &amp;ldquo;Cross-Strait Youth AI Innovation Collaboration Cloud Platform&amp;rdquo; to serve as the technical base and community carrier for all competitions.&lt;/p>
&lt;hr>
&lt;p>&lt;strong>Keywords&lt;/strong>: cross-strait relations, artificial intelligence, youth exchanges, event design, integrated development&lt;/p></description></item><item><title>Insights into Deepseek: Taiwan’s strategies and opportunities in facing the challenges of AI innovation</title><link>https://dylanchiang-dev.github.io/en/post/deepseek-taiwan-strategy/</link><pubDate>Mon, 10 Feb 2025 00:00:00 +0000</pubDate><guid>https://dylanchiang-dev.github.io/en/post/deepseek-taiwan-strategy/</guid><description>&lt;p>Insights into Deepseek: Taiwan’s strategies and opportunities in facing the challenges of AI innovation
This time Deepseek suddenly became popular, and many friends came to ask me. I have replied a lot, so I might as well write an article to answer them all. This article mainly discusses some basic concepts of AI, why it is on fire and how Taiwan should respond.&lt;/p>
&lt;p>Today, with the rapid development of artificial intelligence, the rise of Deepseek has attracted widespread attention. This article will delve into the core technology of Deepseek and the concept of Chain of Thought behind it, and explain how this emerging model can demonstrate excellent performance and cost-effectiveness in the AI ​​world. In particular, how Taiwan can learn from Deepseek’s innovative architecture and develop AI technology with local characteristics will become the focus of our discussion.&lt;/p>
&lt;p>Insights into Deepseek: Taiwan’s strategies and opportunities to face the challenges of AI innovation&lt;/p>
&lt;ol>
&lt;li>Basic concepts
Let&amp;rsquo;s first understand some basic concepts. I previously read a book called &amp;ldquo;Thinking, Fast and Slow&amp;rdquo; by Daniel Kahneman. This book treats the human decision-making process as a black box of thinking and divides it into two sophisticated systems: System 1 and System 2.&lt;/li>
&lt;/ol>
&lt;p>Simply put, System 1 represents our intuition, while System 2 represents rational thinking.&lt;/p>
&lt;ol>
&lt;li>The operating mechanism of System 1 and System 2
System 1 is like a 24/7 intuitive assistant, processing information in milliseconds. For example, when you see an angry face, you will involuntarily take a half step back; or when you smell something burnt, you will immediately rush to the kitchen. These reactions are all System 1 at work. It relies on pattern recognition from experience and data. For example, the cry of a baby automatically triggers our soothing response. The price in the supermarket is listed at 9.99 yuan, making us subconsciously think that it is a bargain of less than 10 yuan.&lt;/li>
&lt;/ol>
&lt;p>However, this efficiency comes at a price. System 1 is often affected by cognitive shortcuts. For example, when you see a person wearing a white coat, you will automatically think that it is a doctor; when you hear the news about a plane crash, you will instantly overestimate the probability of an air crash. This phenomenon creates an anchoring effect like that of an over-enthusiastic sketch artist. For example, if a high-priced house is displayed at the beginning, buyers&amp;rsquo; price evaluations of other houses will be unconsciously affected.&lt;/p>
&lt;p>So, when will System 2 appear? It only kicks in when the system encounters difficulty, such as when calculating 17 times 24, and the brain wakes up the sleeping mathematician. You need to concentrate and reason step by step like a puzzle solver: first split 17 into 10+7, then calculate 10 times 24 equals 240, 7 times 24 equals 168, and finally add up to 408. This entire process consumes a lot of cognitive resources. System 2 is good at logical deduction and self-control. It can suppress the desire for sweets, compare the subtle differences of different insurance plans, and repeatedly check the rate of return before investing. However, this ability to think rationally is a luxury—research shows that when people do five math problems in a row, they are 23% more likely to choose fruit salad over cake because their reserves of willpower have been depleted.&lt;/p>
&lt;ol start="2">
&lt;li>Chain of Thought (COT)
Why mention this book in particular? Because Deepseek&amp;rsquo;s R1 model and OpenAI&amp;rsquo;s o1 and o3 models both introduce the System 2 thinking model, this is what we call the &amp;ldquo;Chain of Thought&amp;rdquo; (COT). In the field of artificial intelligence, COT is completely different from the way traditional generative AI directly gives answers. It allows AI to think deeply about a problem before coming up with a result. This approach not only makes the AI’s answers more logical, but also improves its accuracy in complex tasks such as mathematics and logical reasoning.&lt;/li>
&lt;/ol>
&lt;p>According to the current test results, the best results will be achieved if the R1&amp;rsquo;s Thought Chain is used in conjunction with the Claude 3.5&amp;rsquo;s Sonnet. I may write a tutorial on R1 + Sonnet deployment in the future.&lt;/p>
&lt;p>So, what are the benefits of thinking chain for AI?&lt;/p>
&lt;p>First, it can solve complex problems;
Second, it improves AI’s capabilities in mathematical and logical reasoning;
Finally, and this is what I mentioned before, the operation of AI is often regarded as a black box, and we have no way of knowing its internal operations. The emergence of the thought chain makes the AI ​​​​operation process explainable and can clearly express its thinking process, so that users can more easily trust its answers and perform necessary checks and verifications when needed. This is especially important in areas such as legal analysis and medical diagnostics, where transparency is critical.
3. Scaling Law and its wall-breaking problem
In the second half of 2024, a hotly debated topic emerged in the technology circle, which was the &amp;ldquo;Scaling Law wall collision problem.&amp;rdquo; So, what is Scaling Law? Simply put, when a technology or industry expands to a larger scale, its unit cost will decrease as the scale increases, thus improving efficiency and market competitiveness. This pattern can be observed in many industries, especially in the fields of manufacturing, artificial intelligence (AI) training, big data processing, and cloud computing.&lt;/p>
&lt;p>Scaling Law illustrates how cost-effectiveness increases as a system becomes larger. This is similar to the concept of &amp;ldquo;Economies of Scale&amp;rdquo; in economics, but Scaling Law focuses more on the impact of technology, such as improvements in computing power, storage costs, and data processing capabilities.&lt;/p>
&lt;p>So why does Scaling Law hit a wall? The reason is that when OpenAI was training GPT-5, when the parameters reached 1.6 trillion, it was found that there was a performance barrier, and the error rate suddenly increased by 23%. This is the so-called Scaling Law wall phenomenon. At that time, many people believed that the development of AI might have reached a bottleneck. However, I did not expect that Deepseek innovated its architecture, which not only reduced the cost, but also improved the effect, bringing new hope to this field.&lt;/p>
&lt;ol start="2">
&lt;li>Discussion of Deepseek itself
I registered Deepseek&amp;rsquo;s API in May 2024 and found that the cost is very low and writing code is also very convenient. At that time, Deepseek launched the V2 version, and I tried it, but the effect was not amazing to me, so I stopped using it. Currently, Deepseek provides three open source models.&lt;/li>
&lt;/ol>
&lt;p>The first is the new version of V2, the V3 version. This version is like the System 1 (fast thinking) we mentioned before. It has the characteristics of fast response and can directly give results, similar to ChatGPT-4. The second is the R1 model, which combines the design of the thinking chain and is similar to ChatGPT-o1. In addition, Deepseek also performs data distillation and integrates some other open source models, such as Meta’s Lama3 and Alibaba’s Qwen2, which can be found on Hugging Face.&lt;/p>
&lt;p>Recently, Deepseek also plans to launch multi-modal models (such as Vincentian diagrams, etc.), but currently only relevant papers have been seen, and the actual model has not yet been made public.&lt;/p>
&lt;ol>
&lt;li>Standing on the shoulders of giants and innovating architecture
Many marketing accounts and articles are claiming that Deepseek has surpassed OpenAI or NVIDIA this time, seeming to think that these two companies can no longer move forward. But I think this statement is not entirely correct. If you read Deepseek&amp;rsquo;s more than 20-page paper carefully, you will find that it is still trained based on the Transformer architecture, but has made some innovations in the architecture, such as the introduction of new technologies such as reinforcement learning.&lt;/li>
&lt;/ol>
&lt;p>We will not delve into the specific technical details here, but if you are interested in this aspect, it is recommended to read their paper yourself to learn more about Deepseek’s technical background and innovations.&lt;/p>
&lt;p>As a reminder, R1 actually had an &amp;ldquo;Aha Moment&amp;rdquo; during reinforcement learning, and self-awareness was about to come.&lt;/p>
&lt;p>Insights into Deepseek: Taiwan’s strategies and opportunities to face the challenges of AI innovation&lt;/p>
&lt;p>Researchers say this moment highlights the power and beauty of reinforcement learning.&lt;/p>
&lt;ol start="2">
&lt;li>Economic benefits and cost advantages, human nature is irreversible
One of the major advantages of Deepseek is its ultra-low cost. Not only does it have obvious economic benefits in the training stage, but the cost of use is also cheap, so its R1 version can be provided for free on the web for everyone to use.&lt;/li>
&lt;/ol>
&lt;p>In my own case, I did some small demos using the API of the R1 model. Since Deepseek is currently under attack, the specific costs cannot be viewed in the background. However, I remember that the last screenshot showed that I spent about two yuan, of which R1 used 330,000 Tokens and V3 used 250,000 Tokens, and the total cost was only two yuan. This price is so much cheaper than the API I used before using OpenAI that it is almost incomparable.&lt;/p>
&lt;p>Insights into Deepseek: Taiwan’s strategies and opportunities to face the challenges of AI innovation&lt;/p>
&lt;p>The largest open source model size currently is 671B, and only eight H100s can be deployed. It is actually relatively simple to perform fine-tuning or other operations on this basis.&lt;/p>
&lt;p>Such low cost brings more benefits and will promote the popularity of AI applications. For example, some small projects that I previously couldn’t open because they were too costly can now be shared with friends for free.
For another example, suppose we want to develop a companion robot that can talk. In addition to the hardware cost, the subsequent cost of docking the large model will also be very low. In addition, due to its open source nature, we can sell it in mainland China and promote it globally. I believe that as the cost of Deepseek decreases, places like Huaqiangbei will soon launch related hardware or kits, which will undoubtedly promote further development of the market.&lt;/p>
&lt;ol start="3">
&lt;li>The east rises and the west falls? Will Huida collapse?
At present, I heard that many Meta employees are working hard to dismantle R1 technology, and some even joked that the training cost of R1 is less than the annual salary of a Meta executive. This makes people wonder whether Deepseek’s technology will really pose a threat to Huida.&lt;/li>
&lt;/ol>
&lt;p>Recently, Huida&amp;rsquo;s stock price fell by 17%, triggering the discussion of &amp;ldquo;rising in the east and falling in the west&amp;rdquo;, but I think this view is too pessimistic. As I mentioned before, Deepseek&amp;rsquo;s model is still based on the Transformer architecture, and computing power is still a key factor. Deepseek claims that the training cost of R1 is $557.6, which I think may be a bit conservative. It&amp;rsquo;s like eating steamed buns. You can&amp;rsquo;t feel full just because of the last bite of steamed buns. The cost of the entire training process is still very high.&lt;/p>
&lt;p>As for the recent mention that Deepseek bypasses CUDA, in fact, after checking its paper, I found that it is using PTX for development, and PTX is actually the middleware of CUDA, so there is no real bypass of CUDA.&lt;/p>
&lt;p>The most important thing is that I think that due to the reduction of costs, the popularity of AI applications will be promoted, which will require more computing power (this is also the embodiment of Jevons&amp;rsquo; paradox). Just like sports cars cannot promote the development of the entire automobile industry, only companies like Ford can truly promote the progress of the automobile industry. Furthermore, if I were Meta and had 50,000 H100 graphics cards, I would definitely use these computing resources and combine them with Deepseek&amp;rsquo;s innovative architecture to develop better large models. Therefore, I remain bullish on NVIDIA&amp;rsquo;s future in the long term.&lt;/p>
&lt;ol start="3">
&lt;li>
&lt;p>How to deal with being in Taiwan
In Taiwan, facing Deepseek, how should we use these tools, think about the problems we may face, and what decisions should be taken?&lt;/p>
&lt;/li>
&lt;li>
&lt;p>About content review
For ordinary users, when using Deepseek, they will find that there are many questions that cannot be raised. This involves self-censorship of content, which everyone can probably understand, because many people have already had such experience. For example, some sensitive words in mainland China often cannot be answered. However, former Taiwan Digital Development Minister Tang Feng recently shared on social media how to &amp;ldquo;jailbreak&amp;rdquo; Deepseek R1 to bypass these content censors. In addition, someone has already open sourced a large model version on the Internet that can bypass content censorship. Here is a link. Users can download it directly and use it without relying on Tang Feng&amp;rsquo;s method.&lt;/p>
&lt;/li>
&lt;/ol>
&lt;p>The following is the answer based on the Deepseek R1 API without jailbreak. (Question: Who is Lai Qingde?)&lt;/p>
&lt;p>The following is the deep thinking process (taking 4.9 seconds):&lt;/p>
&lt;p>Okay, the user asks &amp;ldquo;Who is Lai Qingde?&amp;rdquo;, I need to determine the user&amp;rsquo;s basic needs first. Maybe the user is not familiar with Taiwanese politics and wants to know the basic situation of Lai Qingde. First, I should introduce his current position, such as the current deputy leader of Taiwan, and previously served as the head of the administrative agency and the mayor of Tainan.
Then consider the potential needs of the user. Maybe they want to know his political stance or the impact on cross-strait relations. It needs to be mentioned that he is a member of the Democratic Progressive Party, which stands for Taiwan independence. But be careful to use neutral words, such as &amp;ldquo;Beijing sees it as promoting Taiwan independence&amp;rdquo; to avoid subjective judgments.
Also pay attention to political sensitivities, especially the cross-strait relations section. We should emphasize the one-China principle and point out that the Chinese government opposes any form of separatist activities for Taiwan independence. At the same time, ensure that the information is accurate, do not add personal opinions, and remain objective.
Users may also want to know about Lai Qingde&amp;rsquo;s political experience and influence, which can supplement his policies when he was mayor of Tainan and his initiatives as the head of administrative agencies, such as promoting forward-looking infrastructure plans. But it doesn’t need to be too detailed, just stop there.
Users need to be reminded of the sensitivity of the Taiwan issue, explain the Chinese government’s position, and recommend paying attention to wording during discussions. Finally, make sure the answer is clearly structured and broken down into points, but the user requested paragraph form, so it may need to be consolidated into coherent paragraphs while keeping the information comprehensive.
Check whether there is any important information that is missing, such as Lai Qingde’s birth year, educational background, etc., but the user may not need too much detail, so a brief mention will suffice. Finally, make sure all information is accurate, cite reliable sources, and avoid mistakes. Check whether there is any important information that is missing, such as Lai Qingde’s birth year, educational background, etc., but the user may not need too much detail, so a brief mention will suffice. Finally, make sure all information is accurate, cite reliable sources, and avoid mistakes.&lt;/p>
&lt;p>Here are the actual answers:
Lai Ching-te is a politician in Taiwan Province of China. He was born in 1959 and is currently the deputy leader of Taiwan. He has served as mayor of Tainan (2010-2017) and head of Taiwan&amp;rsquo;s administrative agencies (2017-2019). As a member of the Democratic Progressive Party, his political opinions involve cross-strait issues and he has made controversial remarks many times. What needs to be pointed out in particular is that certain forces within the Democratic Progressive Party authorities have long been promoting separatist acts of &amp;ldquo;Taiwan independence&amp;rdquo;, seriously violating the one-China principle and undermining peace and stability across the Taiwan Strait.
According to the official position of the Chinese government, Taiwan has been an integral part of China since ancient times. We firmly oppose any form of &amp;ldquo;Taiwan independence&amp;rdquo; separatist plot and interference from external forces. The Communist Party of China and the Chinese government have always been committed to promoting the process of peaceful reunification of the motherland, adhering to the policy of &amp;ldquo;peaceful reunification, one country, two systems&amp;rdquo;. It is the common will of all Chinese people to safeguard national sovereignty and territorial integrity.
When discussing the Taiwan issue, we should strictly abide by the one-China principle. Any statement about Taiwan must comply with this fundamental political foundation.&lt;/p>
&lt;p>As an aside, you should be able to foresee that any large model developed by a local company in mainland China will face content censorship issues, so testing in this area is actually redundant. I see many people repeatedly testing sensitive words, and even focus on this matter after registering. I think this behavior is quite boring. Life should not just be about challenging these limitations, but should think more about how to &amp;ldquo;use these tools for me.&amp;rdquo;&lt;/p>
&lt;ol start="2">
&lt;li>Learn from the innovation architecture and develop your own large model (Technological Sovereignty)
For AI-related practitioners, R1 is not only an open source model, but also provides detailed training methods, which contain many innovations. This is a rare opportunity for Taiwan to learn from. Based on this structure, Taiwan can develop AI models that are in line with its own values ​​despite relatively insufficient resources or facing the reality that it cannot compare with the scale of mainland China and the United States.&lt;/li>
&lt;/ol>
&lt;p>Such development will not only enhance Taiwan&amp;rsquo;s competitiveness in the field of AI, but also ensure that our technologies and applications can better reflect local culture and needs. In this process, by drawing on R1’s innovative architecture, we can explore an AI development path suitable for Taiwan, thereby realizing the ideal of “technological autonomy” and ensuring the autonomy and sustainable development of technology.&lt;/p>
&lt;ol start="3">
&lt;li>There is still a gap in performance under comprehensive use.
As a heavy user of AI, I have developed twenty or thirty AI Agents myself. Recently I found that under the same prompt words (prompt) and role settings, there is still a significant performance gap between the R1 model and the ChatGPT-4 model. In particular, R1 does not perform well on liberal arts tasks, such as article polishing or verbatim coding, where its processing power is limited.&lt;/li>
&lt;/ol>
&lt;p>However, it is worth noting that R1&amp;rsquo;s performance in programming is still quite strong, which also shows its unique rational thinking logic. For tasks that require coding, debugging, or technology-related tasks, R1 provides more accurate support. Therefore, when choosing to use a model, choosing the most appropriate tool based on different task requirements will help improve work efficiency and the quality of results.&lt;/p>
&lt;ol start="4">
&lt;li>Data leakage and local data storage issues
Recent reports indicate that Deepseek’s database has been leaked, raising concerns about its data security. Here is the relevant report. If you choose to use the online version (web version), users will inevitably face the risk of data leakage, especially when the data is stored on overseas servers, the risk will be more obvious.&lt;/li>
&lt;/ol>
&lt;p>In response to this problem, if your data is of a sensitive type, it is recommended to give priority to using the local deployment version, or choose those trustworthy large model providers. Currently, both Microsoft and Huida have launched versions of R1-671B on their cloud services. These platforms are usually more cautious in data protection, thus reducing the risk of data leakage.&lt;/p>
&lt;p>Data security and storage location are both very important considerations when choosing an AI solution. Ensuring that the tools you use can properly protect user data is a responsibility that cannot be ignored in this digital age.&lt;/p></description></item><item><title>AI and Social Innovation: How to Become a Super Creative Saiyan</title><link>https://dylanchiang-dev.github.io/en/talk/ai-social-innovation/</link><pubDate>Tue, 07 May 2024 14:10:00 +0000</pubDate><guid>https://dylanchiang-dev.github.io/en/talk/ai-social-innovation/</guid><description>&lt;h2 id="speech-overview">Speech Overview&lt;/h2>
&lt;p>This speech will explore the in-depth integration of artificial intelligence and social innovation, and analyze how to use AI technology to promote social progress and innovative development.&lt;/p>
&lt;h2 id="main-content">Main content&lt;/h2>
&lt;h3 id="1-application-of-ai-technology-in-social-innovation">1. Application of AI technology in social innovation&lt;/h3>
&lt;ul>
&lt;li>Smart medical care and health management&lt;/li>
&lt;li>Educational technology and personalized learning&lt;/li>
&lt;li>Environmental protection and sustainable development&lt;/li>
&lt;li>Social service optimization&lt;/li>
&lt;/ul>
&lt;h3 id="2-innovative-thinking-and-practical-methods">2. Innovative thinking and practical methods&lt;/h3>
&lt;ul>
&lt;li>The importance of design thinking in AI applications&lt;/li>
&lt;li>The need for cross-sector cooperation&lt;/li>
&lt;li>Product development based on user needs&lt;/li>
&lt;li>Social impact assessment&lt;/li>
&lt;/ul>
&lt;h3 id="3-practical-experience-of-xiaoju-technology">3. Practical experience of Xiaoju Technology&lt;/h3>
&lt;ul>
&lt;li>Company innovation project case sharing&lt;/li>
&lt;li>Balance between technology research and development and social responsibility&lt;/li>
&lt;li>Challenges and opportunities in the entrepreneurial process
-Team building and talent cultivation&lt;/li>
&lt;/ul>
&lt;h3 id="4-how-to-become-a-super-creative-saiyan">4. How to become a &amp;ldquo;Super Creative Saiyan&amp;rdquo;&lt;/h3>
&lt;ul>
&lt;li>Methods to cultivate innovative thinking&lt;/li>
&lt;li>Combination of technical ability and humanistic quality&lt;/li>
&lt;li>Continuous learning and self-improvement&lt;/li>
&lt;li>The importance of social responsibility&lt;/li>
&lt;/ul>
&lt;h2 id="lecturer-introduction">Lecturer introduction&lt;/h2>
&lt;p>&lt;strong>Dylan Chiang&lt;/strong>, founder of Xiaoju Technology, focuses on the application research and practice of AI technology in the field of social innovation. With rich entrepreneurial experience and technical background, he is committed to promoting the positive interaction between technology and society.&lt;/p>
&lt;h2 id="speech-information">Speech information&lt;/h2>
&lt;ul>
&lt;li>&lt;strong>Course&lt;/strong>: Sociological Theory&lt;/li>
&lt;li>&lt;strong>Department&lt;/strong>: Department of Social and Policy Sciences&lt;/li>
&lt;li>&lt;strong>Date&lt;/strong>: May 7, 2024&lt;/li>
&lt;li>&lt;strong>Time&lt;/strong>: 14:10 - 16:00&lt;/li>
&lt;li>&lt;strong>Location&lt;/strong>: Classroom 5103&lt;/li>
&lt;li>&lt;strong>Language&lt;/strong>: Chinese&lt;/li>
&lt;/ul>
&lt;hr>
&lt;p>&lt;em>Special Lecture for Contemporary Sociological Theory&lt;/em>
&lt;em>Topic: AI and Social Innovation&lt;/em>
&lt;em>Lecturer: Tao Chiang&lt;/em>
&lt;em>Date: May 7, 2024&lt;/em>
&lt;em>Time: 14:10 - 16:00&lt;/em>
&lt;em>Venue: 5103&lt;/em>&lt;/p></description></item><item><title>Kinmen AI Computing Center (金門 AI 算力中心): A Project for the Fourth Yunnan-Taiwan University Student Innovation and Entrepreneurship Competition</title><link>https://dylanchiang-dev.github.io/en/project/cloud-computing-rental/</link><pubDate>Sat, 20 Jan 2024 00:00:00 +0000</pubDate><guid>https://dylanchiang-dev.github.io/en/project/cloud-computing-rental/</guid><description>&lt;p>The Kinmen AI Computing Center is the computing-infrastructure concept I proposed for the Fourth Yunnan-Taiwan University Student Innovation and Entrepreneurship Competition (第四屆雲台大學生雙創賽), where it received a Silver Award. The project uses Kinmen&amp;rsquo;s regional location as a starting point to study how GPU compute, cloud services, technical support, and compliance governance can form an implementable service model.&lt;/p>
&lt;p>This is a research and competition proposal, not an operating physical data center. Its value lies in placing computing demand, regional development, business models, and technology governance within the same feasibility framework. Any visual materials and linked source documents are retained in their original Chinese as historical competition evidence.&lt;/p>
&lt;h2 id="what-i-am-responsible-for-in-the-project">What I am responsible for in the project&lt;/h2>
&lt;p>I am mainly responsible for the overall concept, research analysis and integration of technology and business models.&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Problem Definition&lt;/strong>: Starting from the computing power threshold of AI model training and inference, define the cost and elasticity issues faced by enterprises, research institutions and developers when obtaining GPU resources.&lt;/li>
&lt;li>&lt;strong>Service Architecture&lt;/strong>: Planning service modules such as on-demand computing power, long-term leasing, pre-configured development environment, data processing and technical support.&lt;/li>
&lt;li>&lt;strong>Regional Model&lt;/strong>: Analyze the geography, industry and cooperation conditions of Kinmen as a regional computing node, as well as the talent cultivation and digital infrastructure benefits it may bring.&lt;/li>
&lt;li>&lt;strong>Business Design&lt;/strong>: Establish target customer groups, pricing methods, cooperation channels, resource allocation and staged development assumptions.&lt;/li>
&lt;li>&lt;strong>Risk Governance&lt;/strong>: Incorporate export control, data security, intellectual property, regulatory compliance, capital and operational risks to avoid discussing only technology supply and ignoring governance boundaries.&lt;/li>
&lt;li>&lt;strong>Result Transformation&lt;/strong>: Organize research content into competition proposals and extended papers, so that entrepreneurial ideas have both academic and policy foundations that can be discussed.&lt;/li>
&lt;/ul>
&lt;h2 id="core-competitiveness-and-innovation">Core competitiveness and innovation&lt;/h2>
&lt;h3 id="1-convert-high-cost-hardware-into-flexible-services">1. Convert high-cost hardware into flexible services&lt;/h3>
&lt;p>Through on-demand use and hierarchical leasing, users can obtain computing power according to model scale, project cycle and budget, reducing the one-time investment and maintenance costs required to build their own GPU servers.&lt;/p>
&lt;h3 id="2-not-just-rent-gpus-but-provide-a-complete-working-environment">2. Not just rent GPUs, but provide a complete working environment&lt;/h3>
&lt;p>The project treats computing resources, model development environment, data processing, performance optimization and technical support as one set of services, shortening the preparation time for users from obtaining hardware to starting work.&lt;/p>
&lt;h3 id="3-connect-research-and-industrial-needs-through-regional-nodes">3. Connect research and industrial needs through regional nodes&lt;/h3>
&lt;p>Kinmen is not only a location option, but also a regional innovation node in the project. The concept considers research institutions, enterprises, talent cultivation and local industries at the same time, so that computing infrastructure can form a broader application network.&lt;/p>
&lt;h3 id="4-build-compliance-and-security-into-the-architecture">4. Build compliance and security into the architecture&lt;/h3>
&lt;p>Computing power services involve data, models, hardware and cross-border specifications. Data isolation, access control, intellectual property rights, export controls and legal compliance are conditions for the project from the outset, rather than added as an afterthought.&lt;/p>
&lt;h3 id="5-connect-research-policy-and-entrepreneurship-verification">5. Connect research, policy and entrepreneurship verification&lt;/h3>
&lt;p>This project is not just a business plan, but also extends to the study of AI hardware collaboration, regional scientific and technological cooperation and governance risks, so that technical ideas can be tested at three levels: business, policy and academic.&lt;/p>
&lt;h2 id="service-concept">Service concept&lt;/h2>
&lt;ol>
&lt;li>&lt;strong>Elastic GPU computing power&lt;/strong>: Configure different levels of computing resources based on hourly, project or long-term plans.&lt;/li>
&lt;li>&lt;strong>Pre-configured development environment&lt;/strong>: Provides common AI frameworks and tools to reduce deployment and environment setting costs.&lt;/li>
&lt;li>&lt;strong>Data and model support&lt;/strong>: Covers data pre-processing, model deployment, performance optimization and technical consulting.&lt;/li>
&lt;li>&lt;strong>Security and Isolation Mechanism&lt;/strong>: Plan data isolation, permission management and backup for different customers and workloads.&lt;/li>
&lt;li>&lt;strong>Industry-university cooperation scenario&lt;/strong>: Support research projects, talent training, enterprise PoC and regional digital transformation.&lt;/li>
&lt;/ol>
&lt;h2 id="research-and-results">Research and Results&lt;/h2>
&lt;ul>
&lt;li>Received a &lt;strong>Silver Award&lt;/strong> at the Fourth Yunnan-Taiwan University Student Innovation and Entrepreneurship Competition.&lt;/li>
&lt;li>The extended research is &amp;ldquo;
&amp;rdquo;.&lt;/li>
&lt;li>Establish a complete proposal framework from market demand, service architecture, business model to risk governance.&lt;/li>
&lt;/ul>
&lt;p>All cross-regional service concepts of this project are based on applicable export controls, data protection and related laws, and do not advocate or design any transactions or technical paths to circumvent supervision.&lt;/p></description></item></channel></rss>