I respectfully disagree! There's a lot of opportunity behind keyboard + mouse + screen.
In a way Bytebot is a maximalist bet on the growth and improvement of multi-modal LLMs. I firmly believe that in a short period of time, the token cost will drop, while the capability increases (both dramatically). It's still uncertain, which makes it a great asymmetric bet.
We don't do any sort grounding or image conversion, and we offer a handful of tools. I'll go into more detail in my next post.
Agreed, and the consistency has improved over time. I remember only a 9 months ago struggling to get a browser agent to accurately click on a checkbox. The growth trajectory is what has us excited.
- The agent has a tool to set it's task to 'completed', 'failed', or 'needs_help', with the last one being a option for human in the loop scenarios. Sometimes the agent gets lazy and says it needs help prematurely.
- Additionally, the agent can create subtasks for itself, either to run immediately, or to schedule in the future. Here it again can call that tool a bit too eagerly, filling duplicate subtasks for a task that involves repetitive work.
- Properly handling super long running tasks, that run for 1+ hours. The context window eventually hits it's limit (this will be addressed this week)
Aside from those top of mind issues, there's a whole bunch of scaffolding issues - filesystem permissions, prompt injection security, i/o support, token cost - lot's to improve!
We're still super early, but already these agents are showing flashes of brilliance, and we're gaining more and more conviction that this is the right form factor
$450 renewal fee, $300 annual travel credit, so the card costs $150 per year.
With points valued at $0.015, you need to earn 10,000 points to break even (150/0.015).
You get 3 points for every $1 spent on food/travel, so you need to spend $3,333/year on those categories to break even. Personally I spend way more, so the card is definitely worth it to me.
1. It depends, you're able to toggle whether the image can be viewed once or multiple times.
2. We think viewing a portion of the image makes it interesting. You can treat an image as a scavenger hunt, and hide clues within it. It also discourages screenshots!
You can already enter a drop-off location. It might be that Uber needs to make this feature more pronounced in the UI.
On the confirmation screen, you can tap a little plus button next to the pickup address to enter your drop-off location. It's come in handy for me a few times.