Junior Data Scientist
For the Below JD- Please apply only of you are fluent in Japanese language.
Junior Data Scientist – Generative AI & Japanese Language AI
Location: Remote (Company based in San Ramon, CA)
Role Overview
We are seeking a Junior Data Scientist – Generative AI & Japanese Language AI to support the development and evaluation of AI-powered products using Generative AI, Large Language Models (LLMs), and Natural Language Processing (NLP) technologies.
The role focuses on improving Japanese-language AI experiences by evaluating model outputs, designing prompts, analyzing performance, and supporting AI application development. The ideal candidate will have strong Japanese language expertise combined with foundational knowledge of Data Science, Machine Learning, and Generative AI.
Education
- Bachelor’s or Master’s degree in Computer Science, Data Science, Artificial Intelligence, Mathematics, Statistics, Computational Linguistics, or related fields.
- Relevant coursework/projects in Machine Learning, NLP, AI, or Data Science preferred.
Experience
- 0–2 years of experience.
- Recent graduates with AI/ML coursework, research, internships, or academic projects are encouraged to apply.
Required Skills
- Native-level Japanese proficiency with excellent speaking, reading, writing, and comprehension skills.
- Strong professional English communication skills.
- Ability to evaluate Japanese AI-generated content for accuracy, fluency, tone, context, and cultural relevance.
- Foundational understanding of:
- Data Science and Machine Learning concepts
- Natural Language Processing (NLP)
- Large Language Models (LLMs)
- Generative AI workflows
- Python programming experience with familiarity in data analysis libraries.
- Knowledge of prompt engineering, LLM testing, and AI output evaluation methodologies.
- Strong analytical, problem-solving, documentation, and communication skills.
- Ability and willingness to learn emerging AI technologies.
Preferred Skills
- Experience with Generative AI models/platforms such as OpenAI, Claude, Gemini, or open-source LLMs.
- Exposure to Japanese NLP, localization, translation evaluation, data annotation, or linguistic quality testing.
- Familiarity with:
- Retrieval-Augmented Generation (RAG)
- Embeddings and vector databases
- AI agents/agentic workflows
- PyTorch or TensorFlow
- REST APIs and Git
- Experience creating or evaluating datasets, prompts, benchmarks, or automated AI test cases.