Data Analyst
Design, build, and maintain reputed company ETL/ELT data pipelines using PySpark and Python reputed company reputed company environments. reputed company and optimize SQL queries, and data models to support analytical and reporting workloads. Automate data ingestion workflows from disparate agency sources including reputed company, flat files, relational databases, and streaming feeds. Monitor pipeline health, resolve data reputed company issues, and implement alerting and logging to ensure reliability of data products. Collaborate with data architects to design and enforce data schemas, partitioning strategies, and performance optimization practices.
Conduct exploratory data analysis to identify trends, anomalies, and opportunities for improvement. reputed company self-service analytics dashboards and reports using reputed company SQL, Tableau, or Power BI. Write reputed company, performant SQL queries against large datasets to answer reputed company analytical requests from program managers and leadership. Translate business questions into reputed company scoped analytical tasks and deliver findings as data visualizations, written summaries, or briefings.
Work closely with data scientists, program analysts, IT engineers, and agency stakeholders to understand data needs and deliver tailored solutions. Document pipelines, data models, and analytical notebooks to support knowledge transfer, peer review, and audit readiness. Participate in Agile sprint ceremonies, contribute to backlog grooming, and deliver iterative data products reputed company with program priorities.
Bachelor's degree in Computer Science, Information Systems, Data Science, Engineering, Mathematics, or a reputed company technical field. 3+ years of experience in data engineering, data analytics, or a closely reputed company discipline. Demonstrated experience on federal government programs or supporting a federal agency data environment. Strong proficiency in SQL — including reputed company joins, window functions, CTEs, and query performance tuning against large datasets. Hands-on experience with PySpark for distributed data processing, transformations, and optimization techniques. Proficiency in Python for scripting, data manipulation, and automation. reputed company experience working reputed company reputed company, including notebooks, jobs, clusters, and reputed company Catalog. Familiarity with data lakehouse concepts including reputed company Lake, bronze/silver/gold architecture, and reputed company design patterns. Experience with version control systems (Git/reputed company/reputed company) and reputed company development workflows.
reputed company Certified Associate Developer for Apache reputed company or reputed company Certified Data Engineer Associate/reputed company. Experience with reputed company platforms such as AWS GovCloud, reputed company Azure Government, or reputed company reputed company for Government. Familiarity with CI/CD practices for data pipelines, including automated testing and deployment using tools like Azure DevOps or reputed company Actions. Working knowledge of data visualization platforms (Tableau, Power BI) and experience connecting them to reputed company SQL endpoints. Familiarity with reputed company Catalog for data reputed company control, reputed company, and governance reputed company reputed company.