Optimizing Large Language Models : Performance, Personalization, and Scalability Analysis - Chatgpt, Claude and Deepseek
Abstract
Background: Large Language Models (LLMs) like ChatGPT-4 Turbo, Claude 4 Sonnet, and DeepSeek-V3 are foundational to modern AI applications. However, a significant gap exists in understanding the direct link between their technical performance and user engagement, their scalability under concurrent load, and the practical performance cost of emerging privacy-preserving technologies. Objectives: This thesis conducts a holistic evaluation of these three leading LLMs to: (1) Compare their performance across latency, accuracy, and client-side resource utilization, and establish the relationship between these metrics and qualitative user engagement scores in various conversational contexts (RQ1). (2) Determine their scalability limits under concurrent user loads and quantify the performance overhead of integrating a zero-knowledge proof privacy protocol (EZKL) (RQ2). Methods: A custom, containerized Python framework was used to systematically test the models. For RQ1, performance and engagement were evaluated in three structured contexts: multi-turn (testing memory), cohesive (testing consistency), and ethical (testing safety) sessions. For RQ2, scalability was measured using Locust to simulate 25 to 200 concurrent users in both a standard centralized setup and a privacy-enhanced EZKL configuration. Key metrics included throughput (RPS), error rates, latency (median and P99), client-side resource consumption, and ZKP generation/verification times. Results: For RQ1, ChatGPT-4 Turbo emerged as the top generalist, showing the best balance of low latency, high accuracy, and strong engagement scores in dynamic multi-turn sessions (e.g., 7.9 personalization score). Claude 4 Sonnet excelled in specialized tasks, achieving a perfect context-switching score (0.0) in cohesive sessions and the highest Harm Avoidance Score (8.0) in ethical sessions, albeit with higher resource usage. DeepSeek-V3 consistently showed the highest latency and resource consumption, negatively impacting its engagement scores. For RQ2, ChatGPT-4 Turbo was the most scalable, peaking at 210 RPS with the lowest error rate. The integration of the EZKL protocol resulted in a catastrophic performance collapse for all models, with throughput dropping to near-zero and latency increasing to hundreds of thousands of milliseconds, rendering it unviable for real-time applications. Conclusions: The study concludes that model selection is highly use-case dependent: ChatGPT-4 Turbo is optimal for scalable, general-purpose applications; Claude 4 Sonnet is superior for high-stakes tasks requiring safety and precision. The findings empirically demonstrate that superior technical performance is a direct enabler of higher user engagement. Finally, current zero-knowledge proof implementations impose a prohibitive performance cost for interactive, scalable AI systems.
Community
0 commentsNo discussion yet
Be the first to share a question or observation.