Exciting updates on Project GR00T! We discover a systematic way to scale up robot data, tackling the most painful pain point in robotics. The idea is simple: human collects demonstration on a real robot, and we multiply that data 1000x or more in simulation. Let’s break it down: 1. We use Apple Vision Pro (yes!!) to give the human operator first person control of the humanoid. Vision Pro parses human hand pose and retargets the motion to the robot hand, all in real time. From the human’s point of view, they are immersed in another body like the Avatar. Teleoperation is slow and time-consuming, but we can afford to collect a small amount of data. 2. We use RoboCasa, a generative simulation framework, to multiply the demonstration data by varying the visual appearance and layout of the environment. In Jensen’s keynote video below, the humanoid is now placing the cup in hundreds of kitchens with a huge diversity of textures, furniture, and object placement. We only have 1 physical kitchen at the GEAR Lab in NVIDIA HQ, but we can conjure up infinite ones in simulation. 3. Finally, we apply MimicGen, a technique to multiply the above data even more by varying the *motion* of the robot. MimicGen generates vast number of new action trajectories based on the original human data, and filters out failed ones (e.g. those that drop the cup) to form a much larger dataset. To sum up, given 1 human trajectory with Vision Pro -> RoboCasa produces N (varying visuals) -> MimicGen further augments to NxM (varying motions). This is the way to trade compute for expensive human data by GPU-accelerated simulation. A while ago, I mentioned that teleoperation is fundamentally not scalable, because we are always limited by 24 hrs/robot/day in the world of atoms. Our new GR00T synthetic data pipeline breaks this barrier in the world of bits. Scaling has been so much fun for LLMs, and it's finally our turn to have fun in robotics! We are creating tools to enable everyone in the ecosystem to scale up with us: - RoboCasa: our generative simulation framework (Yuke Zhu). It's fully open-source! Here you go: http://robocasa.ai - MimicGen: our generative action framework (Ajay Mandlekar). The code is open-source for robot arms, but we will have another version for humanoid and 5-finger hands: https://lnkd.in/gsRArQXy - We are building a state-of-the-art Apple Vision Pro -> humanoid robot "Avatar" stack. Xiaolong Wang group’s open-source libraries laid the foundation: https://lnkd.in/gUYye7yt - Watch Jensen's keynote yesterday. He cannot hide his excitement about Project GR00T and robot foundation models! https://lnkd.in/g3hZteCG Finally, GEAR lab is hiring! We want the best roboticists in the world to join us on this moon-landing mission to solve physical AGI: https://lnkd.in/gTancpNK
Data Quality for AI
Explore top LinkedIn content from expert professionals.
-
-
Sure, anybody can call OpenAI APIs to access cutting-edge models, but let’s be real: the true opportunity for businesses isn’t just plugging into those APIs. It’s about leveraging your most unique competitive advantage: your data. Data is the foundation of any successful AI system. Yet, the journey from raw data to actual value has many challenges: 1. Not enough data? Your model can’t be generalized. 2. Poor-quality data? Expect poor-quality results. 3. Nonrepresentative data? Say hello to biased predictions. 4. Too many irrelevant features? You’re adding noise, not value. 5. Not enough diversity? Your model won’t be robust. Garbage in, garbage out. Even the most advanced model is only as good as the data it learns from. For businesses, the opportunity lies in building data pipelines tailored to their unique context — clean, representative, and enriched with meaningful features. This is how you create an AI that’s not just smart, but aligned with your business goals. The frontier isn’t just in using AI. It’s in using AI to transform your data into a moat your competitors can’t cross.
-
*𝑆𝑖𝑔ℎ* Yet again, I hear another company excitedly talking about implementing AI—integrating it, scaling it, “revolutionizing everything”—and yet they gloss over the need for a robust data strategy. It takes all my energy not to pull my hair out as I cringe, listening to the words. But instead of yelling into the void, I’ve learned a better approach: I ask questions. Good ones. The kind that make leaders pause and realize that AI without solid data foundations is just a very expensive experiment. 𝐐𝐮𝐞𝐬𝐭𝐢𝐨𝐧𝐬 𝐥𝐢𝐤𝐞: 1) What percentage of your data is truly usable—normalized, contextualized, indexed, and properly mapped? 2) How much of your data is “dark” (produced but unused), and what’s your plan to leverage it? 3) Do you have a defined data governance and data management framework, or is it mostly ad hoc? 4) What’s your process for ensuring data accuracy, completeness, and relevance for AI models? 5) How scalable is your data infrastructure to support AI at an enterprise level? 6) If AI solutions depend on a continuous flow of clean data, how confident are you that your processes can deliver that over time? This is when the lightbulb flickers. Because here’s the reality: You already produce more data than you know what to do with. And yet, no one is asking whether your data is reliable, clean, and strategically aligned. Oh, and let’s not forget—you’re probably not even collecting the right strategic data yet to unlock AI’s full potential. AI doesn’t live in isolation. It thrives on organized, high-quality data. Your first step to scaling AI shouldn’t be building models—it should be building a foundation: ✅ 𝐃𝐚𝐭𝐚 𝐢𝐧𝐟𝐫𝐚𝐬𝐭𝐫𝐮𝐜𝐭𝐮𝐫𝐞 ✅ 𝐃𝐚𝐭𝐚 𝐠𝐨𝐯𝐞𝐫𝐧𝐚𝐧𝐜𝐞 ✅ 𝐃𝐚𝐭𝐚 𝐦𝐚𝐧𝐚𝐠𝐞𝐦𝐞𝐧𝐭 ✅ And, most importantly, a 𝐝𝐚𝐭𝐚 𝐬𝐭𝐫𝐚𝐭𝐞𝐠𝐲. 𝐒𝐨 𝐛𝐞𝐟𝐨𝐫𝐞 𝐲𝐨𝐮 𝐝𝐢𝐯𝐞 𝐢𝐧𝐭𝐨 𝐀𝐈, 𝐚𝐬𝐤 𝐲𝐨𝐮𝐫𝐬𝐞𝐥𝐟: “If AI is the engine of innovation, do we even have the fuel to power it?” (Trust me, the answer might surprise you.) ******************************************* • Visit www.jeffwinterinsights.com for access to all my content and to stay current on Industry 4.0 and other cool tech trends • Ring the 🔔 for notifications!
-
According to Gartner, AI-ready data will be the biggest area for investment over the next 2-3 years. And if AI-ready data is number one, data quality and governance will always be number two. But why? For anyone following the game, enterprise-ready AI needs more than a flashy model to deliver business value. Your AI will only ever be as good as the first-party data you feed it, and reliability is the single most important characteristic of AI-ready data. Even in the most traditional pipelines, you need a strong governance process to maintain output integrity. But AI is a different beast entirely. Generative responses are still largely a black box for most teams. We know how it works, but not necessarily how an independent output is generated. When you can’t easily see how the sausage gets made, your data quality tooling and governance process matters a whole lot more, because generative garbage is still garbage. Sure, there are plenty of other factors to consider in the suitability of data for AI—fitness, variety, semantic meaning—but all that work is meaningless if the data isn’t trustworthy to begin with. Garbage in always means garbage out—and it doesn’t really matter how the garbage gets made. Your data will never be ready for AI without the right governance and quality practices to support it. If you want to prioritize AI-ready data, start there first.
-
AI is redefining what’s possible in modern agriculture. Have you tried maximizing parsley cultivation? Consider this: 🌱 52 parsley bunches from a single aeroponic tower occupying just 1 m² 🌱 9-week crop cycle (3 weeks from seed to seedling + 6 weeks from transplant to harvest) 🌱 6 seeds per growing port to maximize uniformity and yield Now imagine combining that system with AI. The Power of AI + Aeroponics Agriculture generates enormous amounts of data every day: Temperature Humidity CO₂ concentration Light intensity (PPFD/DLI) Nutrient EC pH Water usage Growth rates Harvest weights AI can analyze millions of data points continuously and make adjustments faster than any human operator. What AI Can Deliver ✅ 15-30% yield improvement Through optimized environmental control and predictive growth models. ✅ Up to 95% less water Aeroponic systems already use dramatically less water than traditional farming. AI helps optimize every misting cycle and nutrient delivery event. ✅ 20-40% reduction in fertilizer waste By dynamically adjusting nutrient concentrations based on plant growth stages. ✅ Early disease detection Computer vision systems can identify plant stress and nutrient deficiencies days before they become visible to the human eye. ✅ More accurate harvest forecasting AI can predict harvest windows with high precision, helping farms reduce waste and meet customer demand. The Economics Scale Fast A facility with: 1,000 towers 52 bunches per tower 52,000 bunches per harvest cycle Even a modest 10% increase in yield means: ➡️ 5,200 additional bunches every cycle ➡️ More revenue without expanding floor space ➡️ Better return on infrastructure investments The Future: Autonomous Farms The next generation of vertical farms will operate as living AI systems: Cameras monitoring every plant Sensors feeding real-time environmental data AI models predicting growth trajectories Automated nutrient and irrigation control Digital twins simulating thousands of optimization scenarios before changes are made This isn't science fiction. The global AI in agriculture market is projected to grow from billions today to tens of billions of dollars over the next decade as growers seek higher yields, lower resource consumption, and greater food security. The future of farming won't be defined by how much land you have. It will be defined by how intelligently you use every square meter. 🌱 52 bunches. 1 square meter. Millions of data points. One intelligent growing system. #AI #Aeroponics via @agrotonomy #VerticalFarming #SmartAgriculture #FoodTech #AgTech #ControlledEnvironmentAgriculture #MachineLearning #Sustainability #Innovation #DigitalTwin #FutureOfFood #PrecisionAgriculture
-
In the world of Generative AI, 𝗥𝗲𝘁𝗿𝗶𝗲𝘃𝗮𝗹-𝗔𝘂𝗴𝗺𝗲𝗻𝘁𝗲𝗱 𝗚𝗲𝗻𝗲𝗿𝗮𝘁𝗶𝗼𝗻 (𝗥𝗔𝗚) is a game-changer. By combining the capabilities of LLMs with domain-specific knowledge retrieval, RAG enables smarter, more relevant AI-driven solutions. But to truly leverage its potential, we must follow some essential 𝗯𝗲𝘀𝘁 𝗽𝗿𝗮𝗰𝘁𝗶𝗰𝗲𝘀: 1️⃣ 𝗦𝘁𝗮𝗿𝘁 𝘄𝗶𝘁𝗵 𝗮 𝗖𝗹𝗲𝗮𝗿 𝗨𝘀𝗲 𝗖𝗮𝘀𝗲 Define your problem statement. Whether it’s building intelligent chatbots, document summarization, or customer support systems, clarity on the goal ensures efficient implementation. 2️⃣ 𝗖𝗵𝗼𝗼𝘀𝗲 𝘁𝗵𝗲 𝗥𝗶𝗴𝗵𝘁 𝗞𝗻𝗼𝘄𝗹𝗲𝗱𝗴𝗲 𝗕𝗮𝘀𝗲 - Ensure your knowledge base is 𝗵𝗶𝗴𝗵-𝗾𝘂𝗮𝗹𝗶𝘁𝘆, 𝘀𝘁𝗿𝘂𝗰𝘁𝘂𝗿𝗲𝗱, 𝗮𝗻𝗱 𝘂𝗽-𝘁𝗼-𝗱𝗮𝘁𝗲. - Use vector embeddings (e.g., pgvector in PostgreSQL) to represent your data for efficient similarity search. 3️⃣ 𝗢𝗽𝘁𝗶𝗺𝗶𝘇𝗲 𝗥𝗲𝘁𝗿𝗶𝗲𝘃𝗮𝗹 𝗠𝗲𝗰𝗵𝗮𝗻𝗶𝘀𝗺𝘀 - Use hybrid search techniques (semantic + keyword search) for better precision. - Tools like 𝗽𝗴𝗔𝗜, 𝗪𝗲𝗮𝘃𝗶𝗮𝘁𝗲, or 𝗣𝗶𝗻𝗲𝗰𝗼𝗻𝗲 can enhance retrieval speed and accuracy. 4️⃣ 𝗙𝗶𝗻𝗲-𝗧𝘂𝗻𝗲 𝗬𝗼𝘂𝗿 𝗟𝗟𝗠 (𝗢𝗽𝘁𝗶𝗼𝗻𝗮𝗹) - If your use case demands it, fine-tune the LLM on your domain-specific data for improved contextual understanding. 5️⃣ 𝗘𝗻𝘀𝘂𝗿𝗲 𝗦𝗰𝗮𝗹𝗮𝗯𝗶𝗹𝗶𝘁𝘆 - Architect your solution to scale. Use caching, indexing, and distributed architectures to handle growing data and user demands. 6️⃣ 𝗠𝗼𝗻𝗶𝘁𝗼𝗿 𝗮𝗻𝗱 𝗜𝘁𝗲𝗿𝗮𝘁𝗲 - Continuously monitor performance using metrics like retrieval accuracy, response time, and user satisfaction. - Incorporate feedback loops to refine your knowledge base and model performance. 7️⃣ 𝗦𝘁𝗮𝘆 𝗦𝗲𝗰𝘂𝗿𝗲 𝗮𝗻𝗱 𝗖𝗼𝗺𝗽𝗹𝗶𝗮𝗻𝘁 - Handle sensitive data responsibly with encryption and access controls. - Ensure compliance with industry standards (e.g., GDPR, HIPAA). With the right practices, you can unlock its full potential to build powerful, domain-specific AI applications. What are your top tips or challenges?
-
I invited 31 researchers to test AI research synthesis by running the exact same prompt. They learned LLM analysis is overhyped, but evaluating it is something you can do yourself. Last month I ran an #AI for #userresearch workshop with Rosenfeld Media. Our first cohort was full of smart, thoughtful researchers (if you participated in the workshop, I hope you’ll tag yourself and weigh in in the comments!). A major limitation of a lot of AI for UXR “thought leadership” right now is that too much of it is anecdotal: researchers run datasets a few times through a commercial tool and decide whether or not the output is good enough based on only a handful of results. But for nondeterministic systems like generative AI, repeated testing under controlled conditions is the only way to know how well they actually work. So that’s what we did in the workshop. Our workshop participants produced a lot of interesting findings about qualitative research synthesis with AI: 1️⃣ LLMs can product vastly different output even with the exact same prompt and data. The number of themes alone ranged from 5 to 18, with a median of about 10.5. 2️⃣ Our AI-generated themes mapped pretty well to human-generated themes, but there were some notable differences. This led to a discussion of whether mapping to human themes is even the right metric to use to evaluate AI synthesis (how are we evaluating whether the human-generated themes were right in the first place?). 3️⃣ The bigger concern for the researchers in the workshop was the lack of supporting evidence for themes. The supporting quotes the LLM provided looked okay superficially, but on closer investigation *every single participant* found examples of data being misquoted or entirely fabricated. One person commented that validating the output was ultimately more work than performing the analysis themselves. Now, I want to acknowledge that this is one dataset, one prompt (although, a carefully vetted one, written by an industry expert), and one model (GPT 4o 2024-11-20). Some researchers claim that GPT 4o is worse for research hallucinations–and perhaps it is–but it is still a heavily utilized model in current off-the-shelf AI research tools (and if you’re using off-the-shelf tools, you won’t always know which models they’re using unless you read a whole lot of fine print). But the point is–I think this is exactly the level at which we should be scrutinizing the output of *all* LLMs in research. AI absolutely has its place in the modern researcher’s toolkit. But until we systematically evaluate its strengths and weaknesses, we're rolling the dice every time we use it. We'll be running a second round of my workshop in June as part of Rosenfeld Media’s Designing with AI conference (ticket prices go up tomorrow; register with code PAINE-DWAI2025 for a discount). Or, to hear about other upcoming workshops and events from me, sign up for my mailing list (links below).
-
🪂 How To Make Your Design System AI-Ready (https://lnkd.in/dtnpy7CM), a practical guide on how to reduce drifts, minimize mistakes, maintain context and improve the quality of AI-generated prototypes — with structured spec files, automated auditing and token layers. Put together by Hardik Pandya from Atlassian. --- 🔹 1. Design Decisions Are Infrastructure AI-generated prototypes often don't deliver consistently decent results because of tiny inconsistencies scattered all across a design system. Often it's decisions made but not documented, hard-coded values never cleaned up, or relying too much on AI making sense of mock-ups or design flows on its own. Unsurprisingly, better AI prototypes come from better data — but also from better human guidance. We shouldn’t assume that AI knows how to choose the right component, and how to design with accessibility in mind. It needs priorities, a clear path on how we make decisions, design principles, examples, do's and don'ts. In fact, we should treat design decisions as infrastructure. That means that every time we make a decision — not just a design decision, but even decision on how actually prioritize our work and how we make decisions around here — it must find a path into the spec file that is then consumed by AI. --- 🔶 2. Three Layers: Spec Files + Token Layer + Audit To ensure quality, we establish design principles, guidelines, rules in a form of “spec files”). It's structured Markdown files that include spacing rules, color choices, component usage guidelines, priorities etc. AI is going to read and reuse that spec file every time it's going to generate a prototype. Because the spec files are text files, it's much more cost-effective, but also much more accurate just because we don't rely on AI recognizing or decoding patterns from mock-ups, but gets specific guidelines instead. In fact, extending code is often a more effective way than generating code from mock-ups. Token layer lists and keeps updated all tokens used throughout the design system. AI always chooses from a closed set of named variables instead of inventing plausible values ad-hoc. An audit script catches what AI gets wrong. It scans the prototype and flags every hard-coded value and flags it if necessary. It can be a regular software doing that, with AI waiting for its feedback to come back. Finally, when a design system ships updates, a sync routine flags which spec files need updating. The goal is to make sure that AI always reads up-to-date, current specs, not the ones written against an outdated version. --- 🔺 3. Examples of AI-Ready Design Systems ⌾ Atlassian: https://lnkd.in/dVsGc3Cp ⌾ Carbon: https://lnkd.in/d4zq4WWb ⌾ CMS Design System: https://lnkd.in/dHHzV3en ⌾ Nordhealth: https://lnkd.in/d8C4j2ZA Yet again, AI can’t magically resolve technical debt or design debt — it needs guidance, decisions, priorities and principles.
-
𝟯𝟴% 𝗯𝗲𝘁𝘁𝗲𝗿 𝗔𝗜 𝗮𝗰𝗰𝘂𝗿𝗮𝗰𝘆. 𝗡𝗼 𝗻𝗲𝘄 𝗺𝗼𝗱𝗲𝗹. 𝗡𝗼 𝗻𝗲𝘄 𝗱𝗮𝘁𝗮. 𝗝𝘂𝘀𝘁 𝗯𝗲𝘁𝘁𝗲𝗿 𝗰𝗼𝗻𝘁𝗲𝘅𝘁. That’s the headline from a controlled NL-to-SQL experiment discussed by Manoj Shanmugasundaram in Metadata Weekly. Across 522 query evaluations, the only variable that changed was context quality — and it made all the difference. Concise, high-signal context (business definitions, SQL patterns, domain rules) drove a 38% accuracy gain. Verbose, catalog-style documentation? Performance dropped and costs rose. More words diluted the signal. The biggest lift wasn't on simple or extreme queries. It was on medium-complexity ones — the joins and aggregations that make up everyday analytics work — where focused context delivered a 𝟮.𝟭𝟱𝘅 𝗶𝗺𝗽𝗿𝗼𝘃𝗲𝗺𝗲𝗻𝘁. The mindset shift: metadata was built for humans to browse. Now it also needs to work for machines to reason. Most teams are still optimizing for readability, not machine usability. If your "talk to data" initiative is stalling, it might not be a model problem. It might be a context problem. Manoj breaks down what machine-usable context actually looks like — and how to get started without rebuilding your stack. Read it in Metadata Weekly 👇
-
Do you think Data Governance: All Show, No Impact? → Polished policies ✓ → Fancy dashboards ✓ → Impressive jargon ✓ But here's the reality check: Most data governance initiatives look great in boardroom presentations yet fail to move the needle where it matters. The numbers don't lie. Poor data quality bleeds organizations dry—$12.9 million annually according to Gartner. Yet those who get governance right see 30% higher ROI by 2026. What's the difference? ❌It's not about the theater of governance. ✅It's about data engineers who embed governance principles directly into solution architectures, making data quality and compliance invisible infrastructure rather than visible overhead. Here’s a 6-step roadmap to build a resilient, secure, and transparent data foundation: 1️⃣ 𝗘𝘀𝘁𝗮𝗯𝗹𝗶𝘀𝗵 𝗥𝗼𝗹𝗲𝘀 & 𝗣𝗼𝗹𝗶𝗰𝗶𝗲𝘀 Define clear ownership, stewardship, and documentation standards. This sets the tone for accountability and consistency across teams. 2️⃣ 𝗔𝗰𝗰𝗲𝘀𝘀 𝗖𝗼𝗻𝘁𝗿𝗼𝗹 & 𝗦𝗲𝗰𝘂𝗿𝗶𝘁𝘆 Implement role-based access, encryption, and audit trails. Stay compliant with GDPR/CCPA and protect sensitive data from misuse. 3️⃣ 𝗗𝗮𝘁𝗮 𝗜𝗻𝘃𝗲𝗻𝘁𝗼𝗿𝘆 & 𝗖𝗹𝗮𝘀𝘀𝗶𝗳𝗶𝗰𝗮𝘁𝗶𝗼𝗻 Catalog all data assets. Tag them by sensitivity, usage, and business domain. Visibility is the first step to control. 4️⃣ 𝗠𝗼𝗻𝗶𝘁𝗼𝗿𝗶𝗻𝗴 & 𝗗𝗮𝘁𝗮 𝗤𝘂𝗮𝗹𝗶𝘁𝘆 𝗙𝗿𝗮𝗺𝗲𝘄𝗼𝗿𝗸 Set up automated checks for freshness, completeness, and accuracy. Use tools like dbt tests, Great Expectations, and Monte Carlo to catch issues early. 5️⃣ 𝗟𝗶𝗻𝗲𝗮𝗴𝗲 & 𝗜𝗺𝗽𝗮𝗰𝘁 𝗔𝗻𝗮𝗹𝘆𝘀𝗶𝘀 Track data flow from source to dashboard. When something breaks, know what’s affected and who needs to be informed. 6️⃣ 𝗦𝗟𝗔 𝗠𝗮𝗻𝗮𝗴𝗲𝗺𝗲𝗻𝘁 & 𝗥𝗲𝗽𝗼𝗿𝘁𝗶𝗻𝗴 Define SLAs for critical pipelines. Build dashboards that report uptime, latency, and failure rates—because business cares about reliability, not tech jargon. With the rising AI innovations, it's important to emphasise the governance aspects data engineers need to implement for robust data management. Do not underestimate the power of Data Quality and Validation by adapting: ↳ Automated data quality checks ↳ Schema validation frameworks ↳ Data lineage tracking ↳ Data quality SLAs ↳ Monitoring & alerting setup While it's equally important to consider the following Data Security & Privacy aspects: ↳ Threat Modeling ↳ Encryption Strategies ↳ Access Control ↳ Privacy by Design ↳ Compliance Expertise Some incredible folks to follow in this area - Chad Sanderson George Firican 🎯 Mark Freeman II Piotr Czarnas Dylan Anderson Who else would you like to add? ▶️ Stay tuned with me (Pooja) for more on Data Engineering. ♻️ Reshare if this resonates with you!
Explore categories
- Hospitality & Tourism
- Productivity
- Finance
- Soft Skills & Emotional Intelligence
- Project Management
- Education
- Technology
- Leadership
- Ecommerce
- User Experience
- Recruitment & HR
- Customer Experience
- Real Estate
- Marketing
- Sales
- Retail & Merchandising
- Science
- Supply Chain Management
- Future Of Work
- Consulting
- Writing
- Economics
- Employee Experience
- Healthcare
- Workplace Trends
- Fundraising
- Networking
- Corporate Social Responsibility
- Negotiation
- Communication
- Engineering
- Career
- Business Strategy
- Change Management
- Organizational Culture
- Design
- Innovation
- Event Planning
- Training & Development