Listing of Claims
U.S. Patent No. 11,887,367 B1 · Using Machine Learning to Train and Use a Model to Perform Automatic Interface Actions Based on Video and Input Datasets · Baker et al., OpenAI Opco LLC · Reissue draft under 37 C.F.R. § 1.173
1.(Original)
A method for training a machine learning model to perform automated actions, comprising:receiving unlabeled digital video data;generating pseudo-labels for the unlabeled digital video data, the generating comprising: receiving labeled digital video data; training a first machine learning model including an inverse dynamics model (IDM) using the labeled digital video data; and generating at least one pseudo-label for the unlabeled digital video data, wherein the at least one pseudo-label is based on a prediction, generated by the IDM, of one or more actions that mimic at least one timestep of the unlabeled digital video data, and the prediction of the one or more actions is generated based on a non-causal combination of past information and future information within the unlabeled digital video data, the past and future information being relative to one or more reference frames within the unlabeled digital video data;adding the at least one pseudo-label to the unlabeled digital video data to form pseudo-labeled digital video data; andfurther training the first machine learning model or a second machine learning model using the pseudo-labeled digital video data to generate at least one additional pseudo-label for the unlabeled digital video.
2.(Currently amended)
The method of claim 1, wherein the IDM or machine learning model is trained to generate one or more predicted actions to be performed via a graphical user-interface overlay rendered atop a live application windowTab ↹without invoking an operating-system input event on the host machine.wherein the overlay is generated by a second neural network distinct from the IDM.and refreshed at a rate exceeding sixty frames per second.