Data Scientist
Data Scientist with Bachelor’s degree in Computer Science, Computer Information Systems, Information Technology, or a combination of education and experience equating to the U.S. equivalent of a Bachelor’s degree in one of the aforementioned subjects.
Job Duties and Responsibilities:
- Collaborate with product owners, data scientists, data engineers, analysts, architects, and business stakeholders to identify high-value data, AI, and analytics opportunities.
- Translate business requirements into analytical approaches, technical specifications, success metrics, acceptance criteria, and implementation plans.
- Collect, profile, analyze, and interpret large structured, semi-structured, and unstructured datasets to uncover trends, patterns, anomalies, and business opportunities.
- Perform exploratory data analysis, statistical analysis, hypothesis testing, correlation analysis, segmentation, and feature engineering.
- Develop, train, tune, and evaluate machine-learning models for regression, classification, clustering, forecasting, recommendation, optimization, and anomaly detection.
- Apply advanced statistical and machine-learning techniques, including ensemble methods, time-series modeling, causal inference, and predictive analytics.
- Select suitable algorithms and evaluate models using precision, recall, F1 score, AUC, RMSE, MAE, explainability, fairness, and business-impact metrics.
- Design experiments and A/B tests to validate hypotheses, measure model effectiveness, and quantify the business impact of data-driven solutions.
- Develop business dashboards, reports, data visualizations, and self-service analytics solutions that communicate findings to technical and nontechnical stakeholders.
- Define key performance indicators, build reusable analytical datasets, and perform ad hoc and root-cause analyses to support strategic and operational decisions.
- Develop end-to-end Retrieval-Augmented Generation pipelines using enterprise data, embeddings, semantic search, hybrid search, reranking, vector databases, and large language models.
- Apply prompt engineering, grounding, citation validation, structured outputs, evaluation frameworks, and guardrails to improve the accuracy and reliability of Generative AI applications.
- Design Agentic AI solutions capable of planning, reasoning, tool selection, workflow execution, reflection, and autonomous task completion.
- Build single-agent, multi-agent, and supervisor-agent workflows using LangGraph, LangChain, LlamaIndex, or comparable Agentic AI frameworks.
- Integrate AI agents with enterprise APIs, databases, vector stores, documents, cloud services, analytics platforms, and business applications.
- Implement function calling, tool calling, workflow routing, agent memory, state management, checkpointing, and Model Context Protocol integrations.
- Develop human-in-the-loop approvals, confidence thresholds, escalation processes, fallback mechanisms, retry policies, and error-recovery controls for agentic workflows.
- Design scalable data models and analytical schemas using MongoDB, relational databases, cloud data warehouses, data lakes.
- Develop and maintain ETL/ELT pipelines for data ingestion, cleansing, transformation, validation, normalization, enrichment, and preparation.
- Build and support batch, streaming, Change Data Capture, and event-driven data pipelines across operational, analytical, and AI systems.
- Implement data-quality rules, metadata management, lineage, observability, schema-evolution handling, and governance controls.
- Develop Python-based data and AI services and integrate machine-learning, Generative AI, and Agentic AI capabilities through APIs and reusable components.
- Package, deploy, version, monitor, and maintain machine-learning models, AI agents, analytical applications, and data pipelines across cloud and enterprise environments.
- Monitor model accuracy, agent behavior, data drift, model drift, pipeline health, application performance, latency, token consumption, security, and infrastructure costs.
- Evaluate AI solutions for hallucinations, bias, prompt injection, sensitive-data exposure, unauthorized tool usage, and regulatory risks while documenting findings and coordinating delivery across cross-functional, onshore, and offshore teams.
Technologies Involved / Skills required for the position:
- Strong programming skills in Python and SQL, including object-oriented programming, data structures, exception handling, debugging, and performance optimization.
- Experience with data science libraries and machine-learning frameworks such as Pandas, NumPy, Scikit-learn, XGBoost, LightGBM, TensorFlow, and PyTorch.
- Strong knowledge of exploratory data analysis, statistical analysis, hypothesis testing, feature engineering, causal inference, experiment design, and A/B testing.
- Experience developing regression, classification, clustering, forecasting, recommendation, optimization, and anomaly-detection models.
- Knowledge of model-evaluation techniques and metrics, including precision, recall, F1 score, AUC, RMSE, MAE, explainability, fairness, and business-impact measurement.
- Experience developing dashboards, reports, and self-service analytics using Power BI, Tableau, Looker, Omni, or comparable business-intelligence platforms.
- Hands-on knowledge of Generative AI, large language models, prompt engineering, embeddings, grounding, structured outputs, tool calling, and guardrails.
- Experience building RAG solutions using LangChain, LangGraph, LlamaIndex, Hugging Face, or comparable document-processing and AI orchestration frameworks.
- Knowledge of Agentic AI systems involving reasoning, planning, reflection, workflow routing, tool execution, memory, state management, and autonomous task completion.
- Experience implementing function calling, tool calling, API integrations, checkpointing, and Model Context Protocol integrations.
- Knowledge of vector databases and search technologies such as MongoDB Atlas Vector Search, Pinecone and Chroma.
- Experience integrating AI models through OpenAI, Anthropic, Google Gemini, Vertex AI, AWS Bedrock, Azure OpenAI, or Hugging Face APIs.
- Knowledge of developing Python-based backend services and REST APIs using FastAPI, Flask, or comparable frameworks.
- Experience with data-processing and distributed-computing technologies, including Pandas, NumPy, PySpark, SQL, and data-validation frameworks.
- Experience with relational and NoSQL databases, including PostgreSQL, MySQL, SQL Server, MongoDB, and DynamoDB.
- Experience with cloud data warehouses, data lakes, or Lakehouse platforms such as BigQuery, Snowflake, Azure Synapse, or Databricks.
- Experience developing ETL/ELT, batch, streaming, Change Data Capture, and event-driven data pipelines.
- Knowledge of workflow-orchestration and messaging technologies such as Apache Airflow, Kafka, Google Cloud Pub/Sub, AWS SQS, SNS, and Event Bridge.
- Knowledge of data-quality rules, metadata management, data lineage, observability, schema evolution, and data-governance controls.
- Knowledge of deploying AI/ML applications using Docker, Kubernetes, and cloud platforms.
- Understanding MLOps and LLMOps practices, including model versioning, experiment tracking, CI/CD, automated testing, controlled releases, rollback, retraining, and production monitoring.
- Knowledge of AI security and responsible AI practices, including hallucination evaluation, bias detection, prompt-injection prevention, sensitive-data protection, explainability, model governance, privacy, and regulatory compliance.
- Excellent analytical, problem-solving, architecture, code-review, documentation, communication, collaboration, and stakeholder-management skills.
Work location is Portland, ME with required travel to client locations throughout USA.
Rite Pros is an equal opportunity employer (EOE).
Please Mail Resumes to:
Rite Pros, Inc.
565 Congress St, Suite # 305
Portland, ME 04101.
Email: resumes@ritepros.com