AI

Runtime and memory constraints in Llama 3.2 edge deployment

Deploying Llama 3.2 3B on smart glasses via ExecuTorch presents challenges like static context length and lack of KV cache. While Meta Ray-Bans hold 80% market share, privacy concerns persist due to a 2.4% PII memorization rate during model inversion attacks.

Runtime and memory constraints in Llama 3.2 edge deployment

The Llama 3.2 3B model runs on modern smartphones using approximately 2GB of VRAM for Q4_K_M quantization. Developers deploy these models on smart glasses using the ExecuTorch runtime. ExecuTorch uses a static graph execution model that locks the context length before deployment. This limitation means a developer cannot increase the context from 64 tokens to 128 tokens during runtime. The runtime also lacks KV cache features, which makes autoregressive generation impractical for long outputs. The 1B and 3B models use pruning and knowledge distillation to reduce size. The 3B model outperforms models like Gemma 2 2.6B and Phi 3.5-mini on tasks like instruction following and summarization. You should use llama.cpp, given your knowledge of local deployment, if your use case demands heavy reasoning or multi-turn dialogue.

Quantization Model Size VRAM Required Speed (tok/s) Hardware Example
Q2_K ~1.3 GB ~2 GB ~170 iPhone 15 Pro
Q4_K_M ~1.9 GB ~3 GB ~145 Mac M1 8GB
Q5_K_M ~2.2 GB ~3.5 GB ~125 Pixel 8 Pro
Q8_0 ~3.3 GB ~4.5 GB ~90 Mac M2 8GB

Machine Perception and Privacy

The distinction between recording and machine perception is the central privacy problem for the smart glasses industry. On newer models, the LED does not illuminate when a user asks the glasses to identify a landmark or plant. Meta says these images go to an AI model for interpretation rather than a user gallery. This allows the device to analyze the world without giving nearby people the visible signal that Meta uses to defend against covert recording. Software detects if a user covers or damages the LED and shuts down the camera on thousands of devices. A consolidated class action in the U.S. District Court for the Northern California includes many individuals who allege Meta glasses captured them even though they neither bought nor wore the devices. Research shows a 2.4% empirical memorization rate for personally identifiable information during model inversion attacks on the Llama 3.2 1B model.

Workers in Kenya reported watching graphic content like sex and bathroom usage to create AI training data. In Illinois, a proposed class action alleges Meta used Facebook and Instagram photographs to develop a facial recognition system known as NameTag. The NYPD warned other departments in January that smart glasses pose a security threat.

Market Competition and Cognitive Gaps

The smart glasses market shows rapid growth, and Meta Ray-Bans currently hold more than 80% of the AI glasses market share. Samsung expects to launch Galaxy Glasses this year, while Google and Apple also plan to release smart glasses. Snap focuses on its Specs AR hardware and AI-driven advertising tools. Alibaba already released Qianwen AI glasses, which achieved considerable sales.

Current wearable AI systems lack awareness of the user’s internal cognitive state. These systems cannot anticipate user needs without access to cognitive load data. The GazeMind framework addresses this by using eye-tracking data for LLM-based reasoning. The framework achieves state-of-the-art performance by outperforming baselines by over 20% across metrics. Can developers bridge the gap between machine vision and human mental effort?