RESPONSIBILITIES:<br>ML Model Deployment & Platform Management<br>• Lead the design, implementation, and ongoing maintenance of scalable ML infrastructure on Databricks, including ML flow for experiment tracking, model registry, and model serving endpoints.<br>• Oversee the development of the ML Ops platform and automated pipelines for deploying, monitoring, and maintaining models within production environments.<br>• Implement robust solutions for model versioning, systematic retraining, and comprehensive artifact management using Databricks Unity Catalog for ML governance.<br>• Design and manage Databricks Feature Store for consistent feature engineering across training and inference pipelines.<br>Generative AI & LLM Operations<br>• Architect and implement Retrieval-Augmented Generation (RAG) systems for document Q&A, enabling business teams to query fund documents, investor letters, and market research.<br>• Design, deploy, and manage vector database solutions (Databricks Vector Search, Pinecone, or similar) for semantic search and retrieval across enterprise documents.<br>• Lead LLM fine-tuning and customization initiatives, training models like Claude or open-source alternatives with CIM proprietary data while ensuring data privacy and compliance.<br>• Develop and optimize document processing pipelines including PDF parsing, chunking strategies, and embedding generation for RAG applications.<br>• Implement prompt engineering best practices and LLM evaluation frameworks to ensure output quality, relevance, and factual accuracy.<br>• Build guardrails and safety measures for GenAI applications, including hallucination detection, output validation, and source attribution.<br>Automation & CI/CD Pipelines<br>• Design and implement extensive automation across the ML workflow, covering model training, testing, validation, and deployment using Databricks Workflows and Asset Bundles.<br>• Set up robust CI/CD pipelines for both traditional ML models and GenAI applications, leveraging GitHub Actions, Azure DevOps, or similar tools.<br>• Automate complex data and model workflows utilizing orchestration tools such as Airflow, Prefect, or Databricks Workflows.
<p><strong>Machine Learning Engineer</strong></p><p><br></p><p><strong>Company Overview</strong></p><p>Based in Los Angeles, California, the company specializes in transforming complex, multi-source data into actionable insights through machine learning, knowledge graph technologies, and advanced analytics. This is an opportunity to work on mission-critical applications in a highly collaborative environment focused on innovation, scalability, and operational excellence.</p><p><br></p><p><strong>Role Summary</strong></p><p>The Machine Learning Engineer will play a critical role in designing, training, deploying, and optimizing machine learning models that operate on large-scale temporal, geospatial, relational, and unstructured datasets. This position requires an experienced engineer who can independently own the full machine learning lifecycle, from dataset development and model architecture selection to deployment, monitoring, and continuous improvement. The ideal candidate brings broad expertise across computer vision, natural language processing (NLP), geospatial analytics, MLOps, and large-scale production machine learning environments.</p><p><br></p><p><strong>Key Responsibilities</strong></p><ul><li>Design, train, evaluate, deploy, and optimize machine learning models across multiple production use cases.</li><li>Build predictive solutions for anomaly detection, forecasting, entity resolution, relationship prediction, risk assessment, and operational decision support.</li><li>Partner with data engineering teams to develop high-quality training datasets from structured, unstructured, temporal, relational, and geospatial data sources.</li><li>Design model architectures and select appropriate algorithms based on business objectives, data characteristics, and operational requirements.</li><li>Develop and maintain machine learning pipelines spanning data preparation, feature engineering, training, evaluation, deployment, and monitoring.</li><li>Build scalable solutions that leverage graph-based and knowledge graph-driven data architectures.</li><li>Develop models utilizing computer vision, NLP, geospatial analytics, and predictive modeling techniques.</li><li>Establish rigorous evaluation frameworks, baselines, performance metrics, and validation methodologies.</li><li>Design experiments that mitigate data leakage, model drift, bias, and changing data distributions.</li><li>Implement monitoring, observability, alerting, retraining, rollback, and model governance processes.</li><li>Maintain reproducible datasets, model artifacts, evaluation results, and deployment workflows.</li><li>Collaborate with distributed engineering teams to deliver reliable and scalable machine learning capabilities.</li><li>Improve model calibration, confidence scoring, uncertainty estimation, and explainability.</li><li>Contribute to technical architecture, machine learning standards, and long-term platform strategy.</li></ul><p><strong>Additional Details</strong></p><ul><li>Fully onsite 5 days a week</li><li>Highly collaborative environment with strong emphasis on machine learning, knowledge graphs, and data-driven decision support</li><li>Opportunity to influence technical direction, machine learning standards, and model lifecycle practices across multiple initiatives</li><li>Candidates must be authorized to work in the United States and satisfy applicable regulatory employment requirements</li></ul>