Remote Freedom

GenAI Data Engineer

Tiger Analytics

Remote · United States United StatesFull-timetoday

See how well you match this role

Upload your résumé — we'll score your fit and check off the skills you already have. Free.

Requirements

  • Snowflake expertise including unstructured data capabilitiesMust
  • Google Cloud Platform (GCP) experienceMust
  • DataikuMust
  • Python for data processingMust

What you'll do

  • Design and implement robust data pipelines
  • Ingest, process, and store unstructured data formats at scale
  • ETL/ELT process design and implementation
  • SQL query optimization and tuning
  • AI and machine learning tool integration (OCR, NLP, Document AI)
  • Cloud-native data infrastructure design
  • Data integration between GCP and Snowflake
  • Petabyte-scale data environment optimization
  • Collaboration with data scientists and analysts
  • Support data-driven decision making across the organization

Keywords

BigQueryCloud StorageDataflow

About the role

Tiger Analytics is a fast-growing advanced analytics consulting firm. Our consultants bring deep expertise in Data Science, Machine Learning and AI. We are the trusted analytics partner for several Fortune 100 companies, enabling them to generate business value from data. Our business value and leadership has been recognized by various market research firms, including Forrester and Gartner. We are looking for top-notch talent as we continue to build the best analytics global consulting team in the world. We are seeking an experienced Data Engineer with expertise in Dataiku to join our data team. As a Data Engineer, you will be responsible for designing, building, and maintaining data pipelines, data integration processes, and data infrastructure. You will collaborate closely with data scientists, analysts, and other stakeholders to ensure efficient data flow and support data-driven decision making across the organization. Requirements • Design and implement robust data pipelines that ingest, process, and store unstructured data formats at scale within Snowflake and GCP . • Leverage Snowflake’s unstructured data capabilities (Directory Tables, Scoped URLs, Snowpark) to make "dark data" queryable and actionable. • Build and maintain cloud-native ETL/ELT processes using BigQuery, Cloud Storage, and Dataflow, ensuring seamless integration between GCP and Snowflake. • Instead of just using LLMs, you will integrate AI tools (OCR, NLP entities, Document AI) into the engineering flow to transform unstructured blobs into structured insights. • Tune complex SQL queries and Python-based processing jobs to handle petabyte-scale environments efficiently. Benefits Significant career development opportunities exist as the company grows. The position offers a unique opportunity to be part of a small, challenging, and entrepreneurial environment, with a high degree of individual responsibility. Originally posted on Himalayas

Sourced from Himalayas. Confirm the details and apply on the employer's site.