The Real Problem: Productive Data Scientists vs. Leaky Data
If you're running ML workloads on sensitive data, you already know the tension. Data scientists need fast, flexible access to tools. Security teams need to guarantee that data never leaves the perimeter. Traditional answers — air-gapped on-prem boxes, proctored VDI sessions — collapse the moment your team scales past a handful of people.
A real-world case from the AWS Architecture Blog illustrates this perfectly. A fintech organization (referred to as iBusiness) was burning $40+ per user per month on dedicated virtual desktops, each requiring a 2-day provisioning SLA and constant patching. When the team grew, that model broke.
This article walks through the three-layer security architecture they landed on — and that you can adapt for your own ML platform.
TL;DR: Managed browser + URL allowlisting + VPC endpoints + no-NAT SageMaker VPC = 80% cost reduction and airtight exfiltration controls.
![]()
The Three-Layer Defense Model
Layer 1 — Lock the Front Door: WorkSpaces Secure Browser
Amazon WorkSpaces Secure Browser is a managed, Chromium-based browser that runs inside your own VPC. No local install, no persistent session, no user-controlled extensions.
Key controls enforced at this layer:
# WorkSpaces Secure Browser configuration
browser_policy:
disable_file_download: true
disable_file_upload: true
disable_clipboard: true
disable_printing: true
network:
vpc: data-science-vpc
subnet: private-subnet-1
outbound: nat-gateway
nat_elastic_ip: "203.0.113.42"
iam_condition:
# Only allow requests originating from the NAT gateway's Elastic IP
aws:SourceIp: "203.0.113.42"
The IAM side is where most teams get lazy. Don't just allow sagemaker:* — gate it on aws:SourceIp matching the NAT gateway's EIP. That single condition eliminates the "data scientist logs in from home laptop" bypass.
Layer 2 — Constrain the Browser and Cross-Account Movement
A locked-down browser is useless if users can curl data to a random S3 bucket in another AWS account. Two controls close that gap:
1. Strict URL allowlisting
# Permitted URL patterns (everything else is blocked)
*.aws.amazon.com
*.sagemaker.aws.amazon.com
*.studio.sagemaker.aws.amazon.com
# Explicitly blocked
*mail.google.com*
*dropbox.com*
*drive.google.com*
*transfer.sh*
2. VPC endpoints for Console + IAM Identity Center
Instead of letting traffic hit the public console.aws.amazon.com, you route it privately:
# Private Route 53 hosted zone overrides
console.aws.amazon.com -> vpce-0abc123.console.us-east-1.vpce.amazonaws.com
*.console.aws.amazon.com -> vpce-0abc123.console.us-east-1.vpce.amazonaws.com
signon.aws.amazon.com -> vpce-0def456.sso.us-east-1.vpce.amazonaws.com
Combine this with Route 53 Resolver DNS Firewall in the SageMaker VPC. Any DNS query to a non-approved domain gets NXDOMAIN. This kills DNS tunneling exfiltration, which is one of the sneakiest channels attackers (and careless insiders) use.
Finally, add an IAM policy that denies any action whose target resource lives in another AWS account:
{
"Effect": "Deny",
"Action": "*",
"Resource": "*",
"Condition": {
"StringNotEquals": {
"aws:ResourceAccount": "${aws:PrincipalAccount}"
}
}
}
Layer 3 — Harden the SageMaker AI Environment Itself
SageMaker Studio gives users a terminal and IDE. That's a feature and an exfiltration risk. The fix is to remove internet egress entirely from the SageMaker VPC:
# Terraform sketch: SageMaker VPC with no internet egress
resource "aws_vpc" "sagemaker" {
cidr_block = "10.20.0.0/16"
enable_dns_hostnames = true
enable_dns_support = true
# NOTE: no aws_internet_gateway, no NAT gateway attached
}
resource "aws_vpc_endpoint" "s3" {
vpc_id = aws_vpc.sagemaker.id
service_name = "com.amazonaws.us-east-1.s3"
vpc_endpoint_type = "Interface"
policy = data.aws_iam_policy_document.s3_restrict.json
}
Key rules:
- No NAT gateway, no internet route in the SageMaker VPC.
- VPC endpoints for every AWS service SageMaker needs (S3, ECR, CloudWatch, KMS, STS, etc.).
- Endpoint policies scoped to your org's account IDs — not
"Resource": "*". For example, allows3:PutObjectonly to specific bucket ARNs.
This is the layer that makes the architecture actually trustworthy. If someone finds a shell escape, there's nowhere to send the bytes.

What This Architecture Does Not Solve
A three-layer model is strong, but not bulletproof. Be honest with your stakeholders about these edges:
| Risk | Mitigation Status | Notes |
|---|---|---|
| Malicious insider with legitimate allowlisted access | ⚠️ Partial | Audit logs (CloudTrail + VPC Flow Logs) are your only backstop |
| Steganographic exfiltration via model weights | ❌ Not covered | If a user can PutObject to a bucket, they can encode data in tensors |
| Compromised AWS service credentials | ⚠️ Partial | Endpoint policies help, but a leaked role session is dangerous |
| Side-channel via timing or error messages | ❌ Not covered | Requires application-level hardening |
| Human error in IAM policy drift | ⚠️ Partial | Use SCPs + IAM Access Analyzer, not just inline policies |
Practical warning: The most common failure mode isn't a clever attacker — it's an over-permissive VPC endpoint policy that someone added during a Friday deploy. Treat endpoint policies as code, review them in PRs, and alert on any CreateVpcEndpoint or ModifyVpcEndpoint API call.
Also, keep in mind that aws:SourceIp conditions break the moment anyone uses a proxy or VPN. Pair them with aws:SourceVpce conditions wherever possible — they're harder to spoof.
The Cost Story (Because Finance Will Ask)
| Metric | Before (VDI) | After (WorkSpaces Secure Browser) |
|---|---|---|
| Cost per user / month | $40+ | ~$7 |
| Provisioning SLA | 2 days | Minutes (automatic) |
| Ongoing desktop maintenance | High (patching, image updates) | Near-zero (managed) |
| Scalability ceiling | ~dozens of users | Hundreds+ |
That's roughly an 80% cost reduction with stronger security posture. This is the rare case where compliance and FinOps actually agree.
![]()
Where to Go From Here
If you're building something similar, sequence the work like this:
- Audit first. Map every path data currently takes out of your ML environment. You'll find at least one you forgot.
- Start with Layer 3. Removing internet egress from the SageMaker VPC is the highest-leverage change and requires the least organizational buy-in.
- Then Layer 1. Migrate users to WorkSpaces Secure Browser. Expect pushback; measure productivity before and after.
- Layer 2 last. URL allowlisting and DNS Firewall are the most operationally disruptive — you'll need a real change-request process.
Further Reading
- React Compiler v1.0 Is Here: A Deep Dive into Automatic Memoization — if you're also shipping frontend tooling alongside your ML platform, this is worth a read.
- React's New Chapter: The React Foundation Launches Under the Linux Foundation — governance matters for long-lived infra, and this is a good case study.
The bottom line: you don't need air-gapped hardware to protect sensitive ML data. You need disciplined network boundaries, scoped IAM, and a managed browser that removes the human from the attack surface. Do those three things, and your security team will stop blocking your roadmap.