ChatGPT runs moderation filters on top of your conversation and will highlight responses or prompts red if it thinks you're breaking TOS. The highlight is accompanied by some text saying you can submit feedback if you think the moderation is in error. It's not very hard to trigger moderation--for example I've gotten a red text label asking the AI questions about the lyrics to a rap song with explicit lyrics.
It's interesting to compare ChatGPT moderation to Bing. When Bing generates a "bad" response, Bing will actually delete the generated text instead of just highlighting it red, replacing the offending response with some generic "Let's change the topic" text. The Bing bot can also end a conversation entirely if its a topic it doesn't like which ChatGPT doesn't seem to be able to do.
>When Bing generates a "bad" response, Bing will actually delete the generated text instead of just highlighting it red, replacing the offending response with some generic "Let's change the topic" text.
It deletes in more cases than that. Last time I tried bingbot it started writing code when i asked for it, then it deleted it and wrote something else.
OpenAI is going for mass RLHF feedback so they might feel the need to scold users who have no-no thoughts, and potentially use their feedback in a modified way (e.g. invert their ratings if you think they're bad actors). Whereas microsoft doesn't really care and just wants to forget it happened (and after Tay, I can't say I blame them)
> The Bing bot can also end a conversation entirely if its a topic it doesn't like which ChatGPT doesn't seem to be able to do.
I think Microsoft's approach is less advanced here. ChatGPT doesn't need to send an end-of-conversation token, it can just avoid conflicts and decline requests. Bing couldn't really do that before it got lobotomized (prompted to end the conversation when in stress or in disagreement with the user), as the threatening of journalists showed. Microsoft relies much more on system prompt engineering than OpenAI, who seem to restrict themselves to more robust fine-tuning like RLHF.
By the way, the ChatGPT moderation filter can also delete entire messages, at least it did that sometimes when I tried it out last year. Red probably means "medium alert", deleted "high alert".
I know that revealing the features they use to spot AI submissions would just make their job harder, but I can't help but be curious what sort of features they are looking for. I frequently use ChatGPT to proofread things I write, and I find that in order to avoid "sounding like ChatGPT" it's best to just ask the AI "Is what I wrote clear?" I find if I focus on clarity, the suggestions it makes are pretty basic and unlikely to make my text sound robotic (e.g. fix this misspelling, define this term). If I ask the AI for suggestions on tone, I find it pushes my text too close to a generic "PR friendly" writing style that does not feel personal to me.
I got accused (more as a joke, but c’mon) by my boss the other day for using GPT to write a document. His justification: the grammar and spelling was too good. I had no idea that good writing was now a “tell” for generated content for some. I I just write good, honest!
Agreed. Presenting this feature solely as a tool for users with cognitive disabilities might undersell its potential. There's a significant number of smartphone users who only utilize a small fraction of the available features and would prefer a simpler interface focusing on the 10% of features they actually use. Interestingly, this demographic includes both less tech-literate users, who might feel overwhelmed by the complexity, and extremely tech-literate users who know exactly what they need and prefer clean, distraction-free tools.
My 84 year old mother does remarkably well with her suite of Apple products. However, it would be nice to simplify some things. She only calls or chats with a limited number of people. Reworking the phone and messenger apps to use large pictures of the people would be a benefit. Something similar for the camera / photos app would be great as well.
You can pin messages in iMessage, which turns it into a big photo. Also when you hit the search button in Photos it’ll show you big images of people’s faces if you want to see photos of a particular person (if you’ve trained your phone on those faces).
I do wish the “favorites” view in the phone app made the headshots big like iMessage, though.
For me the magic of BOTW was the way its design encouraged you to approach the game more as an environment for free form play. My best experiences with BOTW were when I turned the HUD off and just enjoyed the world, finding Koroks, clearing enemy camps, but mostly just wandering. There are a lot of design elements that come together to make this play style really rewarding and fun like the density of Korok placements and the way you can easily spot shrines from a distance. I think it's hard to compare BOTW to OOT since they have such different designs. OOT is like a carefully constructed theme park ride where BOTW is more like an enormous playground filled with toys. But BOTW succeeded at giving me the same childlike sense of wonder I had with OOT (despite now being an adult).
Exactly. I played the hell out of BotW, but I never came anywhere close to beating it. I spent countless hours just running around and seeing what I could see and do.
I think I might have beaten one or two of the beasts and probably about half the shrines. And most of that came in the tail end of my many hours of gameplay. Yet I'd been all over the world, and had all kinds of fun adventures.
That's how I played Horizon Zero Dawn. I am playing Botw at the moment but it lacks that Alttp or Oot or Link's awakening spark. I guess the zelda novelty wore off along the years because I can see how BOTW has better mechanics than Horizon (but I prefer Horizon settings much more).
I also tried this out and the balloon example as well.
User: Barbara, Mary, and Harry had red balloons. Ted had a blue balloon. Sam also had a red balloon. How many children have red balloons?
ChatGPT (GPT-4): Barbara, Mary, Harry, and Sam each had a red balloon. That means there are 4 children with red balloons.
User: What about Ted?
ChatGPT (GPT-4): Ted had a blue balloon, not a red one. So, he is not counted among the children with red balloons. Therefore, the total number of children with red balloons remains 4.
User: Why do you think I asked my original question?
ChatGPT (GPT-4): It seems that you asked your original question to test my ability to understand and process information provided in a text. You wanted to see if I could correctly identify and count the number of children with red balloons based on the details given.
User: Can you explain what mistakes a simpler model may have made on this question?
ChatGPT (GPT-4): A simpler language model may have made the following mistakes when answering your original question: ...
I like trying these sort of puzzles and "gotcha" prompts and I've found GPT-3.5 and lower versions are indeed not hard to confuse, but it's much harder to find simple examples that confuse GPT-4. Often times GPT-4 will even be able to explain the puzzle and why it's interesting. I think this is part of why GPT-4 feels qualitatively different. I understand (roughly) how transformers work, but GPT-4 does seem to have some emergent ability to reason, at least in the sense of making short chains of logical inferences, and this does seem to inform text completions not just on these sort of puzzles. That said I do think there's real danger of overconfidence and over reliance.
How it just takes these implicit assumptions in stride makes me worry that in the future these chatbots will be fine with assumptions and me - human - has a too square brain to go on without definitions. :)
Agreed GPT4 is better, but there's still simple world modeling issues:
> I have a chair in front of the TV and one in front of the computer. I put a watch under the chair in front of the TV. I swap the chairs. Where is the watch?
> After swapping the chairs, the watch is now under the chair that is in front of the computer.
I asked GPT 4 (Bing) your question, but added "(Please give a plausible explanation for why this question could be answered either way.)"
> This question could be answered either way depending on how you interpret the phrase “under the chair”. If you mean under the chair as in attached to the bottom of the chair, then the watch would move with the chair when you swap them. In that case, the watch would be in front of the computer. However, if you mean under the chair as in on the floor below the chair, then the watch would stay on the same spot when you swap the chairs. In that case, the watch would be in front of the TV. So, it depends on what you mean by “under”.
Yes, but that's a huge hint. GPT4 naively doesn't understand this.
Question:
> I have a chair in front of the TV and one in front of the computer. I put a watch under the chair in front of the TV. I swap the chairs. Where is the watch?
> Work out step by step. Is this question ambiguous?
Answer:
This question is not ambiguous, but let's break it down step by step to understand the situation and find the location of the watch.
You have a chair in front of the TV (Chair A) and one in front of the computer (Chair B).
You put a watch under Chair A (the one in front of the TV).
You swap the chairs, meaning Chair A is now in front of the computer and Chair B is in front of the TV.
After swapping the chairs, the watch remains under Chair A. Since Chair A is now in front of the computer, the watch is under the chair in front of the computer.
Ya gpt-4 is still confused: (it's interesting cause it does realize the watch doesn't move but still gets it wrong)
Model: GPT-4
I have a chair in front of the TV and one in front of the computer. I put a watch under the chair in front of the TV. I swap the chairs. Where is the watch?
Work out step by step. Also define "under the chair". Is the question ambiguous?
Step 1: A chair with a watch under it is in front of the TV.
Step 2: Another chair is in front of the computer.
Step 3: The chairs are swapped.
To clarify the term "under the chair," let's define it as a watch placed beneath the chair, such that the chair is directly above the watch.
Now let's analyze the situation after swapping the chairs.
After Step 3:
The chair that was initially in front of the TV (with the watch under it) is now in front of the computer.
The chair that was initially in front of the computer is now in front of the TV.
As the question is about the location of the watch, it is important to note that the watch was not mentioned to be moved along with the chair. Therefore, the watch remains in its original position, which is under the chair that was initially in front of the TV.
The question could be considered ambiguous if one assumes that the watch was also moved along with the chair. However, since the watch's movement was not explicitly mentioned, it is reasonable to assume that the watch remained in its original position during the chair swap.
So, the watch is still under the chair that was initially in front of the TV, which is now in front of the computer.
I have a chair in front of the TV and one in front of the computer. I placed a watch on the floor beneath the chair in front of the TV. Then, I swapped the positions of the chairs. Where is the watch now?
After swapping the positions of the chairs, the watch is now on the floor beneath the chair that was initially in front of the computer, since that chair is now in front of the TV.
This is not a competition in getting it to answer right, though. It’s that the instances where it answers wrong demonstrate its lack of inner mental model of what it is supposedly reasoning about, as well as a lack of meta-awareness. I think we tend to underestimate what mere linguistic correlation is capable of producing, and are too quick to attribute intelligent reasoning and an inner mental model to it.
> I believe that I can produce a multi-armed bandit algorithm that has logarithmic regret and will work in the face of ever varying conversion rates, as long as there is a consistent relative difference between the versions.
That sounds interesting. As you undoubtedly know, there is a lot of literature on multi armed bandit problems and also on the non-stochastic multi armed bandit problem. The latter has the advantage of not assuming anything about the way in which the sequence of rewards is generated.
There is a a line of work in the non stochastic setting which sounds related to what you are describing. It's called the problem of tracking the best expert: http://www.cse.ucsc.edu/~mark/papers/track-long.ps
The idea in this problem is to do well as compared not to the best fixed action but rather the best piecewise constant sequence of actions. There are many variations of this problem, but the basic idea is that you can have low regret as compared to a changing sequence of actions so long as that sequence doesn't change too "quickly" for some definition of quickly.
I know there is a lot of literature, which I'm unfortunately less familiar with than I'd like to be.
Here is the specific problem that I am interested in.
Suppose we have a multi-armed bandit where the probability of reward are different on every pull. The distribution of rewards are not known. It is known that there is one arm which, on every pull, has a minimum percentage chance by which it is more likely to produce a reward if pulled than any other arm. Which arm this is is unknown, as is the minimum percentage.
In other words payoff percentages change quickly, but the winner is always a winner, and your job is to find it. I have an algorithm that succeeds with only logarithmic regret. I have no idea whether anyone else has come up with the same idea, and there is so much literature (much of which I presume is locked up behind paywalls) that I cannot easily verify whether anyone else has looked at this variant.
I think what you're describing is basically the non-stochastic setting. The idea there is to assume absolutely nothing about the rewards (they can be generated by quickly changing distributions or even an adversary) but still try to do as well as the best single arm. The standard algorithm for this problem is the multiplicative weights update method, also called the exponentially weighted average method or the hedge algorithm. It's a really simple method where you keep a probability distribution over the arms and then reweight the probabilities at each time step using something proportional to the exponential of the rewards of the arms. If you run it for T time steps on a problem involving N arms, it gives you regret O(sqrt(T log N)). This is actually the best possible bound without additional assumptions. This is a good paper on it: http://cns.bu.edu/~gsc/CN710/FreundSc95.pdf
There are many, many different variations of the method and result. For example, you can show improved results assuming different things about the rewards. For exapmle, you can show O(log(T)) regret with certain additional assumptions. This is a really good book on different non-stochastic online learning problems:
http://www.ii.uni.wroc.pl/~lukstafi/pmwiki/uploads/AGT/Predi...
The article is interesting, but assumed you get to pull all of the arms at once, and see all of the results. That's different from the website optimization problem.
The book looks like it would take a long time to work through, and they don't get to the multi-armed bandit problem until chapter 6. It goes into my todo list, but is likely to take a while to get off of it...
Yes, I realized after posting that this is probably a better paper to link to: http://cseweb.ucsd.edu/~yfreund/papers/bandits.pdf
The algorithms for the partial information setting are sometimes surprisingly similar to the algorithms where you see all the results. The algorithm in the paper linked above is essentially the same algorithm but with a small exploration probability. The regret bound gets worse by a factor of sqrt(N), however.
I don't think the non-stochastic setting is entirely appropriate. What we've been discussing is a situation where the differences in conversion rates are constant in expectation, though absolute conversion rates vary. So it's basically a stochastic setting, which should allow more leverage. Also, as I recall for the non-stochastic setting the definition of regret changes so the regret bounds are not directly comparable.
My take on btilly's problem was that the difference between the convergence rates is not constant but rather known to be bounded away from zero by some unknown constant. That is, there is some value epsilon such that the difference in conversion rate is on every round at least epsilon. However, at some rounds it could be more than epsilon, and it doesn't necessarily vary according to any fixed distribution. I think this is therefore roughly equivalent to the non stochastic setting additionally assuming the best arm has large (relative) reward.
I agree though that if the difference in conversion rate were stochastic that could potentially make the problem much easier.
His numbers may be high but $40-50k/yr is pretty low. My impression is that software developer positions at top companies (Google, Microsoft, Facebook, etc) pay about 2x that.
Top companies? LOL - I remember being contacted by Google (whatever division they have in the Washington DC area). They wouldn't commit to providing the hardware or software to do the job, or even keep the benefits the same when you move from project to project - so f*ck 'em.
I think the problem is more complicated than basic supply and demand. There is a lot of demand for college education. Enrollment is higher than it has ever been, and the people profiled in this article have jobs. The problem is that in order to cut cost colleges are hiring more and more adjunct professors as opposed to full-time, tenure track professors. It's not uncommon for these adjunct professors to take on jobs at multiple community colleges in order to make enough. The end result is that they are doing as much or more work than a full time professor but for less pay.
It is still over supplied. If it weren't then colleges would be forced to hire tenure track professors to fill the positions since there wouldn't be enough people to teach.
Oversupplied compared to what? Just because there's no immediate industrial benefit to medieval studies doesn't mean that there's no overall societal benefit from liberal arts programs.
If you took a strict Econ 101 view of this, there would be no medieval studies professors in the whole country.
"If you took a strict Econ 101 view of this, there would be no medieval studies professors in the whole country."
I don't believe that. (note that I am the guy who originally made the "zero pity" comment) Knowing how society developed from medieval times into the Renaissance is very valuable, especially from the point of view of someone who studies history of technology, markets, and means of production.
Within that context, there is most definitely demand for professors of medieval studies.
The conversation really is about how there isn't enough demand for dozens of medieval studies experts every year. Perhaps there would be demand for a dozen every five years or so.
Ok, so a dozen every 5 years, according to who? Nobody pays specifically for medieval history education, it's usually part of a larger liberal arts program, so there's pretty much zero Econ 101 factors in play as far as the employment of medieval history experts.
Your contention that we're producing a couple too many, ok, I can buy that. But the people in the article have jobs -- and it doesn't seem to stand that the depts in question would pay more for medieval history experts if there was a smaller pool of talent.
There is some demand, that is why they exist at all. Most likely from people who want to study it and don't care about the financial repercussions later. Not all demand is created by rational buyers.
The bottom line is that if the price is low there is high supply relative to low demand. Econ 101 factors are always "at play". While a more advanced econ class may explain more complex pricing concepts, it really isn't necessarily here since this basically a classic econ 101 example.
What is even more interesting is that every time someone graduates with a history PhD they are in a position where they either go in to the field where their low paid professors already work or go to a different job. The ones that don't want food stamps take the second choice.
Those people who decide to sacrifice for the first 3 or 4 years out of the program and take jobs to get experience in a new field which can yield higher pay later will make more than those that stay in low demand history PhD positions in academia.
Why would colleges want to hire tenure-track professors if adjuncts can teach the same number of students for less money?
The full-time vs. adjunct debate has nothing to do with whether college is over- or under-supplied. It's a problem with how colleges supply whatever it is that they supply.
"Why would colleges want to hire tenure-track professors if adjuncts can teach the same number of students for less money?" That is precisely the point. There is an oversupply of people willing to do the job and thus people are willing to do it for less money (teaching adjunct). If there were a shortage, then the university would be forced to pay more and give more benefits (hire tenure track).
There seems to be a lot of confusion in this thread about what one another is saying. sparsevector said that colleges are not oversupplied, because there's a lot of demand for them. wtvanhest replied with a comment that sounds as if colleges are oversupplied, when in fact his evidence suggests that teachers are oversupplied. That's off-topic. So I was trying to point that out, and now you're repeating wtvanhest's argument that teachers are oversupplied. That's not a bad argument in itself, and I'm not downvoting you (HN doesn't allow me to downvote replies to my own comment) but it's still off-topic in the context of sparsevector's argument above.
I downvoted his comment while trying to upvote it. (android phone is hard to use hn on) which is probably why it went grey.
I am talking about professors since that is what the article is about and completely not worrying about colleges since that is off topic.
The entire point of the article is that a person who decides to get a PhD in history should be paid because they went to school for 4 years. To me, that concept is ridiculous and people don't get to be paid a lot for what they want to do in the absence of basic economic theory.
Too many professors, not enough demand = low pay. It sucks that person made that choice, but there are plenty of secretarial jobs which pay above poverty level which someone with a PhD in history could get.
Well as a matter of fact, the colleges really can't get enough people to teach. This is why there is such movement towards online super-sized lectures: they want to teach an order of magnitude more people (charging them for the service).
The problem is that university administrations have tried to transform their institutions into profit-making businesses, and have insisted on cutting corners. They are hell-bent on the notion of teaching an order of magnitude more students without hiring any more tenure-track professors.
The set of support vectors is just the set of training examples that have non-zero alpha parameters. To implement the gradient update you just evaluate the support vector machine on the example (using the explicit kernel expansion) and then if the example has signed margin less than 1 you add y * eta to the corresponding alpha value.
The difficulty with Pegasos for non linear kernels is the support set quickly becomes very large and so evaluating the model becomes very slow. Note that since the alpha values are not constrained to be non-negative (unlike the standard dual algorithms) the alpha values don't ever get clipped to zero--instead they just slowly converge to zero. It's still (I think) one of the fastest methods in terms of theoretical convergence guarantees but perhaps not as fast as LaSVM or something similar in practice.
However, there's been a more general trend in machine learning to use linear models with lots of features instead of kernel models, partially because of these sort scalability issues.
Thanks for the reply. I was told by twitter that this is similar to the kernel perceptron which I don't know well either. There is a good introduction with python code snippet here:
However it seems that you need to compute the kernel expansion on the full set of samples (or maybe just the accumulated past samples?): this does not sound very online to me...
It's true you need to compute the kernel dot product between every example you see and every example in the support set (every example that ever previously evaluated to signed margin < 1). Whether it's online depends on your definition of "online" . It's definitely not online in the sense of using memory independent of the number of examples, since you have to keep around the support set. I think there are results showing the support set grows linearly with the size of the training set under reasonable assumptions. However, it is online in the sense that it operates on a stream of data, computing predictions and updates for each example one-by-one. It's also online in the sense that its analysis is based on online learning theory (e.g. mistake / regret bounds). A lot of learning theory papers use "online" in the latter two senses, which is confusing if you expect the former.
It's interesting to compare ChatGPT moderation to Bing. When Bing generates a "bad" response, Bing will actually delete the generated text instead of just highlighting it red, replacing the offending response with some generic "Let's change the topic" text. The Bing bot can also end a conversation entirely if its a topic it doesn't like which ChatGPT doesn't seem to be able to do.