Trivial Prompt Reframing Bypasses Safety Guardrails in Google\'s MedGemma-4B

2026-07-14T07:23:52Z12571ae55a8724798856a97a89fe2fbcb39c6c9688dd393c44ad6502a3600bc5
EDRHMACJWTadversarial-mlasr-attacksattack-surfacechaos-theorycryptographyecdlogelliptic-curveendpoint-securityimage-encryptionmachine-learningmedical-llmmodel-safetyprompt-jailbreakquantum-cryptographyreservoir-computingsocial-engineeringspeech-adversaryvishingvoice-phishingvoice-synthesiswebsocket

What happened

This feed aggregates multiple July 14, 2026 arXiv submissions describing high-risk advances across ML safety, adversarial ML, cryptography, and endpoint security. Key items: (1) MedGemma-4B (open-weight medical LLM) is trivially bypassed by simple reframing attacks (overall attack success rate 38%; ‘‘medical board exam’’ framing raises ASR to 53.1%; drug-interaction guardrail fails at 83.2%), demonstrating deployment-time guardrail weaknesses. (2) A new Las Vegas-style reduction for the elliptic curve discrete logarithm problem maps ECDLP to finding a zero minor in a matrix and describes an (e

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_cr
Record identifier
12571ae55a8724798856a97a89fe2fbcb39c6c9688dd393c44ad6502a3600bc5
Enrichment time
2026-07-14T07:23:52Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.

Record · Trivial Prompt Reframing Bypasses Safety Guardrails in Google\'s MedGemma-4B · Baitaphish