Executive Summary
The Tellabs AWS Glue integration project was designed to establish a secure, scalable, and automated data integration pipeline between Oracle NetSuite and AWS. Leveraging AWS Glue, Amazon S3, and orchestration services such as Step Functions and EventBridge, the solution ensures efficient data extraction, transformation, and storage. The Objective was to build and implement automated ETL pipelines to ensure key reporting metrics are improved in terms of speed and accuracy.
A strong emphasis was placed on security, governance, monitoring, and automation, ensuring that all workloads align with AWS best practices. The implementation provides enhanced visibility, operational resilience, and cost optimization for data processing workflows.
About Client
Tellabs is a technology-driven organization requiring secure and scalable cloud infrastructure to support its data integration and analytics workloads. The customer needed a robust AWS-based solution to handle data ingestion from Oracle NetSuite while maintaining strict governance, security, and compliance standards.

Objectives
- Establish a secure AWS account governance model
- Enable seamless integration with Oracle NetSuite
- Implement automated ETL pipelines using AWS Glue
- Ensure high availability and performance monitoring
- Maintain secure access control and identity management
- Enable cost visibility and optimization for Glue jobs
Scope of Engagement
- AWS Account Setup & Governance
- Identity & Access Management (IAM / SSO)
- AWS Glue ETL Pipeline Development
- Monitoring & Logging Framework
- Network Security Implementation (VPC, SG, NACL)
- Deployment Automation using CloudFormation
- Cost Analysis & Optimization
- Runbook Creation & Operational Support
Architecture Overview
The architecture consists of:
- Source System: Oracle NetSuite
- Processing Layer: AWS Glue
- Storage Layer: Amazon S3
- Orchestration: AWS Step Functions & EventBridge
- Secrets Management: AWS Secrets Manager
- Monitoring: Amazon CloudWatch & CloudTrail
This architecture ensures scalability, automation, and secure data processing pipelines.
Solution Overview – To implement automated ETL pipelines using AWS Glue
The solution integrates Oracle NetSuite with AWS using Glue-based ETL pipelines. Data is extracted, transformed, and stored in S3, enabling downstream analytics and reporting.
Key capabilities:
- Automated ETL workflows
- Event-driven execution using EventBridge
- Secure credential handling via Secrets Manager
- Centralized logging and monitoring
- Scalable serverless architecture
Security & Governance
AWS Account Governance SOP
- Root account restricted to initial setup only
- Mandatory MFA enabled on root account
- Use of corporate email & contact details
- CloudTrail enabled across all regions
- Logs stored in secure S3 bucket with deletion protection
Identity & Access Management
- Principle of Least Privilege
- Use of IAM roles & temporary credentials
- Integration with AWS SSO / Active Directory
- No wildcard permissions in policies
- Individual access accountability
Network Security (VPC)
- Security Groups for controlled traffic
- Subnet-level control via Network ACLs
- Restricted DB access to application layer only
- Controlled internet access pathways
Data Security
- Encryption using AWS KMS CMKs
- SSL/TLS for data in transit
- Key rotation enabled
- Fine-grained access control policies

Implement automated ETL pipelines using AWS Glue and ensure to manage Security and Access control
Monitoring & Logging
- CloudWatch Dashboards for:
- Glue job performance
- Step Functions execution
- EventBridge events
- CloudTrail for audit logging
- Logs exported to Amazon S3
- SNS alerts for:
- Job failures
- Latency spikes
- Credential access issues

Implement automated ETL pipelines using AWS Glue and ensure monitoring and Observability by automating data pipelines
Implementation
Infrastructure as Code
- AWS CloudFormation used for:
- Infrastructure provisioning
- Security configurations
- Networking setup
CI/CD Automation
- Integrated with AWS CodeBuild
- Automated deployment pipeline
- No manual console changes
Glue Implementation – automated ETL pipelines using AWS Glue
- ETL jobs configured with optimized DPUs
- Logging enabled for all executions
- Job performance tracked via CloudWatch
Runbook & Troubleshooting
Routine Monitoring Tasks
- Monitor AWS Lambda execution metrics
- Check API Gateway latency & security logs
- Review Aurora DB performance metrics
Troubleshooting Scenarios
Lambda Errors
- Analyze CloudWatch logs
- Adjust timeouts/configurations
API Gateway Latency
- Identify backend bottlenecks
- Optimize integrations
Aurora DB Issues
- Optimize queries & indexing
- Resolve connection bottlenecks
High Latency Scenario (OPE-001)
- Trigger incident response
- Identify root cause
- Apply performance optimizations
Deployment Readiness Checklist
Testing
- Unit Testing
- Integration Testing
- System Testing
- User Acceptance Testing (UAT)
Automation
- CI/CD pipelines
- Automated testing frameworks
- Security & code quality scans
Documentation
- Deployment guide
- Rollback strategy
- Configuration records
Validation
- Pre & post deployment checks
- Deployment checklist completion
- Evidence of successful rollout
Cost Optimization & Performance Tuning for running and implementing automated ETL pipelines using AWS Glue
- Glue pricing based on DPU usage & runtime
- Cost analysis using:
- CloudWatch logs
- AWS Cost Explorer
- Optimization strategies:
- Reduce over-allocated DPUs
- Use Glue Python shell jobs for smaller workloads
- Enable job bookmarks to avoid reprocessing
- Automation via:
- boto3 scripts
- Athena-based reporting
Challenges & Resolutions
| Challenge | Resolution |
| Secure access management | Implemented IAM roles & SSO |
| Monitoring complexity | Centralized CloudWatch dashboards |
| Cost visibility | Implemented tagging & cost analysis |
| Data security compliance | Used KMS CMKs & encryption |
| Deployment consistency | Adopted CloudFormation IaC |
Project Completion
- Successfully deployed automated ETL pipelines
- Established secure AWS governance model
- Enabled real-time monitoring and alerting
- Improved performance and reduced operational risks
- Delivered scalable and maintainable architecture
Note
This implementation follows AWS best practices for:
- Security
- Reliability
- Performance Efficiency
- Cost Optimization
- Operational Excellence

Implement automated ETL pipelines using AWS Glue high-level system setup.
Reference Links
https://docs.aws.amazon.com/prescriptive-guidance/latest/patterns/build-an-etl-service-pipeline-to-load-data-incrementally-from-amazon-s3-to-amazon-redshift-using-aws-glue.html
Read more about Glue
https://aws.amazon.com/glue/
Read more here about our services
AWS Glue Services
- https://www.peritossolutions.com/services/aws-glue-serverless-data-integration/
AWS consulting Services
- https://www.peritossolutions.com/aws-consulting
About Client
1Place is a technology-driven organization focused on leveraging data to enhance decision-making, reporting, and operational intelligence. The company manages a growing data ecosystem comprising multiple data sources, analytics tools, and business applications. AWS offers glue as a fully managed etl service
To address scalability challenges and improve data automation, 1Place collaborated with Peritos Solutions, an AWS Advanced Consulting Partner, to design and implement a serverless data platform using AWS Glue. The goal was to replace manual data workflows with a secure, automated ETL framework that ensures accuracy, consistency, and governance.
Project Background – Data Modernization through AWS Glue ETL service
1Place data operations were previously dependent on traditional ETL tools that lacked automation and flexibility. These manual processes caused data silos, inconsistent data quality, and delayed reporting.
Peritos Solutions proposed an AWS Glue-based serverless ETL framework to transform 1Place data landscape. The solution automated schema detection, data cataloging, and transformation pipelines while ensuring end-to-end visibility through AWS CloudWatch and centralized governance.
This transformation allowed 1Place to manage large-scale data workloads with minimal operational overhead while maintaining compliance, traceability, and cost efficiency.

Objectives of the Engagement
- Establish a serverless, automated data integration framework using AWS Glue ETL service
- Replace legacy ETL pipelines with scalable and efficient Glue jobs
- Implement centralized metadata management using AWS Glue Data Catalog.
- Enable cross-service data integration with S3, RDS, and Redshift.
- Ensure security, governance, and compliance through IAM, encryption, and monitoring.
Scope & Requirements
Scope
The project’s scope included the design, deployment, and optimization of AWS Glue components to enable seamless data flow across 1Place AWS environment.
Key deliverables included:
- AWS Glue Crawlers for automated schema detection.
- Glue Jobs for data transformation and enrichment.
- Glue Workflows for end-to-end orchestration.
- Centralized Data Catalog for metadata governance.
- CloudWatch monitoring and alerting integration.
Requirements
Functional:
- Automated ETL pipeline creation and scheduling.
- Dynamic schema detection and updates.
- Integration with Amazon Redshift and Athena for analytics.Non-Functional:
- Serverless architecture for scalability.
- Secure access control with IAM.
- Centralized monitoring and auditing.
- Cost efficiency and fault tolerance.

AWS Glue ETL service Implementation support for Pipeline Automation
Solution Overview -AWS Glue ETL service
Business Problem Addressed
1Place existing ETL infrastructure was time-intensive and not scalable. Manual intervention led to delays, higher costs, and inconsistent data quality.
Proposed AWS Glue-Based Solution
Peritos Solutions implemented an end-to-end AWS Glue solution integrating multiple data sources and automating data transformation and cataloging. Using Glue Crawlers, Jobs, and Workflows, the entire ETL process became event-driven, reducing human intervention and operational latency.
Key Benefits
- Serverless data integration with zero infrastructure management.
- Automated data discovery and schema management.
- Faster and more reliable data transformation pipelines.
- Improved data governance through centralized metadata.
- Seamless analytics enablement through Athena and QuickSight integration.
Implementation -AWS Glue ETL service
Architecture Overview
The architecture consisted of:
- Data Sources: S3, RDS, and on-premises data via secure connectors.
- ETL Layer: AWS Glue Crawlers, Jobs, and Workflows.
- Data Catalog: Centralized schema and metadata management.
- Analytics Layer: Athena and QuickSight for visualization.
- Monitoring & Logging: CloudWatch for logs, metrics, and alerts.
Technology Stack
- AWS Services: Glue, S3, RDS, Redshift, CloudWatch, Lambda, Secrets Manager, IAM.
- Security: KMS encryption, MFA-enabled IAM roles, and cross-account logging.
- Automation: CI/CD pipelines with AWS CodePipeline and CodeBuild.
AWS Glue Components Implemented
- Glue Crawlers: Automated schema discovery for S3 and RDS datasets.
- Glue Jobs: ETL scripts built using PySpark to clean, normalize, and enrich data.
- Glue Workflows: Orchestration for dependency-based execution.
- Data Catalog: Managed metadata, table schemas, and data lineage.
- Triggers: Event-driven execution using CloudWatch and EventBridge.
Security and Compliance
- IAM policies applied with least privilege.
- Glue roles restricted to authorized services only.
- KMS encryption applied for data at rest and in transit.
- CloudTrail enabled for audit trails and compliance verification.
Runbook and Troubleshooting Scenarios
Routine Operational Tasks
- Daily monitoring of Glue job metrics and DPU utilization.
- Reviewing failed job logs and rerunning based on SLA thresholds.
- Verifying Data Catalog updates and schema integrity.
- Checking Glue job triggers and workflow dependencies.
Common Troubleshooting Scenarios
- Job Failures Due to Schema Drift: Re-run Glue Crawler, refresh Data Catalog, and update ETL script mapping.
- Performance Degradation: Tune Spark configurations and increase DPU allocation.
- Connection Errors: Validate IAM permissions, VPC configurations, and network paths.
- Data Quality Issues: Use Glue dynamic frames and AWS Deequ for validation.

AWS Glue ETL service Implementation support
Deployment Readiness Checklist
Testing
- Unit, integration, and system testing of all Glue jobs.
- Validation of schema mapping, data accuracy, and job success rates.
Automation
- CI/CD pipelines integrated for job versioning and automated deployment.
- Security scans embedded in build pipelines.
Documentation
- Deployment runbook, rollback plan, and configuration details maintained.
Monitoring & Validation
- Glue job metrics and alerts verified in CloudWatch.
- Post-deployment validation ensured job stability.
Evidence: Deployment logs, Glue job screenshots, and automation reports attached to project documentation.
Cost Optimization and Performance Tuning
- Used Glue 3.0 for faster job performance and improved scaling.
- Optimized DPU allocation and job parallelism.
- Leveraged job bookmarks for incremental data loads.
- Enabled data partitioning in S3 for query efficiency.
- Monitored spend through AWS Cost Explorer and adjusted scheduling.
Challenges and Resolutions
| Challenge | Resolution |
| Schema evolution from multiple data sources | Automated schema updates via Glue Crawlers |
| Long-running ETL jobs | Spark job optimization and dynamic partitioning |
| Data duplication in catalogs | Automated Data Catalog cleanup and versioning |
| Integration with legacy databases | Implemented secure JDBC connections and Glue connections |
| Monitoring job failures | Integrated CloudWatch alerts with email/SNS notifications |
Project Completion – AWS Glue ETL service
Deliverables
- AWS Glue Data Catalog, Crawlers, Jobs, and Workflows.
- CloudWatch dashboards for Glue performance monitoring.
- Operational Runbook and Troubleshooting Guide.
- Deployment Readiness Checklist and Evidence Reports.
- CI/CD pipelines for Glue job automation.
Support
Post-implementation support for two months, including 20 hours/month of operational support, bug fixes, and performance optimization.
Next Phase
- Integrate with AWS Lake Formation for enhanced data governance.
- Implement data lineage tracking and metadata versioning.
- Expand to real-time streaming ETL using AWS Glue and Kinesis.
- Develop monitoring dashboards using QuickSight for Glue job analytics.
- Conduct quarterly optimization reviews for cost and performance improvements.

AWS Glue ETL service Implementation support
Reference Links AWS Glue ETL service
https://docs.aws.amazon.com/prescriptive-guidance/latest/patterns/build-an-etl-service-pipeline-to-load-data-incrementally-from-amazon-s3-to-amazon-redshift-using-aws-glue.html
Read more about Glue
https://aws.amazon.com/glue/
Read more here about our services
AWS Glue Services
- https://www.peritossolutions.com/services/aws-glue-serverless-data-integration/
AWS consulting Services
- https://www.peritossolutions.com/aws-consulting









